Software

Most of what I build at FamilyTreeDNA runs behind a product rather than shipping as a package. What follows is the part you can actually open: live tools, runnable tutorials, open repositories, and the documentation that specifies how the methods work.

FamilyTreeDNA tools

Production features I designed or co-designed. Several are public; the rest need a free FamilyTreeDNA account. Every Discover link below works with any valid haplogroup substituted into the URL.

Mitotree and the mitochondrial haplotree

The largest human phylogeny yet described: 53,588 haplogroups from 331,221 complete mitogenomes. I led the project and designed the search, QC, and dating framework behind it.
How it works.

Age estimates

A pocket watch resting on printed DNA sequence

The TMRCA method used across both haplotrees. It uses a relaxed clock, so each branch carries its own rate, and accounts for variable sequencing coverage instead of assuming it away. Calibrated and validated against documented genealogies.
How it works.

Globetrekker

Globetrekker map tracing a paternal lineage from Africa into Europe

Reconstructs the migration route of a paternal lineage across tens of thousands of years, drawn on a map of the world as it existed at the time, with sea level, ice sheets, and ocean currents rebuilt per date. Y-DNA only for now.
y-dna/A/globetrekker
How it works.

Ancestral paths and connections

Ancestral path
Every haplogroup between the root and yours. mtdna/U5a2b2a3/path
Compare
Where two haplogroups share a common ancestor, and when. mtdna/U5a2b2a3/compare
Notable connections
Historical figures on the same branch. mtdna/U5a2b2a3/notable
Ancient connections
Archaeological samples sharing your lineage. mtdna/U5a2b2a3/ancient
Match time tree
Your own matches placed on the dated tree. Requires a test and login; logged-out visitors see a demo. mtdna/U5a2b2a3/matches

Substitute any Y haplogroup and swap mtdna for y-dna to see the paternal equivalents.
For example: y-dna/A/classic

myOrigins 3.0 and the Chromosome Painter

Chromosome painting showing population ancestry along each chromosome

Estimates global ancestry proportions and per-chromosome local ancestry jointly rather than in sequence, which keeps the segments and the percentages consistent and allowed the reference panel to grow from 24 populations to 90. I invented the method and authored the white paper.
How it works.

Family Finder Matching 5.0

Family Finder matching and relationship estimation

Autosomal matching and relationship estimation that classifies whether a match pair is endogamous before predicting the relationship, which removes a systematic bias that inflates kinship estimates in founder populations. Co-authored.
How it works.

Archaic Origins

Neanderthal skull

Neanderthal, Denisovan, and ghost-archaic ancestry as calibrated percentages and per-chromosome segments. Built and validated; release planned on Discover.
How it works.

Open source

Mitotree

Mitotree repository

The Mitotree release: tree files, haplogroup definitions, and the material needed to assign haplogroups against it. Distributed through HaploGrep v3 plugins, so an external lab can use it without changing its existing workflow.
github.com/genebygene/mitotree

Geo-genetic triangulation

Geo-genetic triangulation repository

Infers where someone’s ancestors lived by finding long DNA segments they share with people who have documented genealogies, then correcting for how heavily each region is represented in the reference data. Developed for the Beethoven genome study; the version used in that paper is archived at Dryad.
github.com/paulmaier/geo-genetic-triangulation

fasta2genotype.py

fasta2genotype repository

Converts RAD sequence data into any of several genotype formats, using whole-haplotype information rather than discarding all but one SNP per locus. Written for my dissertation, where standard pipelines were throwing away most of the signal, and recently updated for current Python.
github.com/paulmaier/fasta2genotype

Everything else is at github.com/paulmaier.

Tutorials

Complete, runnable walkthroughs in R, written for graduate students. Each one loads its own page.

Bioclimatic variables mapped across the Yosemite toad rangePopulation genomics: outlier loci and future adaptationExtracting environmental data, finding outliers with PCAdapt and Bayenv2, then mapping current and future local adaptation with gradient forest and generalized dissimilarity modeling.Simulated migration paths across a three-dimensional landscapeAnalysis of population structure: microsatellites vs. ddRADseqComparing the power of microsatellites and ddRADseq to resolve population structure, to see what each marker type can and cannot tell you.Three-dimensional interpolated surface of genetic diversityInterpolating genetic diversity in 3DTurning point estimates of diversity into a continuous surface across a landscape, and rendering it as a rotatable three-dimensional plot.Network of California herpetologists and their academic advisersPedigree of California herpetologistsAn interactive academic genealogy of 288 herpetologists and their advisers, colored by year, sized by number of connections, shaped by study taxon.

Documentation and data

  • myOrigins 3.0 white paper — full method description: phasing, segment classification, the conditional random field, phase correction, and how the 90-population reference panel was built.
  • Family Finder Matching 5.0 white paper — the matching algorithm, endogamy classification, and relationship estimation.
  • Anaxyrus hybrid panel — the diagnostic SNP panel and methods for identifying western toad and Yosemite toad hybrids.

Archived data and analysis scripts

Much of my production code is proprietary. Where that is the case I have linked the method description instead, which is detailed enough to reimplement. If you want to discuss any of it, or you are stuck reproducing something, get in touch.