Software
Most of what I build at FamilyTreeDNA runs behind a product rather than shipping as a package. What follows is the part you can actually open: live tools, runnable tutorials, open repositories, and the documentation that specifies how the methods work.
FamilyTreeDNA tools
Production features I designed or co-designed. Several are public; the rest need a free FamilyTreeDNA account. Every Discover link below works with any valid haplogroup substituted into the URL.
Mitotree and the mitochondrial haplotree
The largest human phylogeny yet described: 53,588 haplogroups from 331,221 complete mitogenomes. I led the project and designed the search, QC, and dating framework behind it.
How it works.
Age estimates
The TMRCA method used across both haplotrees. It uses a relaxed clock, so each branch carries its own rate, and accounts for variable sequencing coverage instead of assuming it away. Calibrated and validated against documented genealogies.
How it works.
Globetrekker
Reconstructs the migration route of a paternal lineage across tens of thousands of years, drawn on a map of the world as it existed at the time, with sea level, ice sheets, and ocean currents rebuilt per date. Y-DNA only for now.
y-dna/A/globetrekker
How it works.
Ancestral paths and connections
- Ancestral path
- Every haplogroup between the root and yours. mtdna/U5a2b2a3/path
- Compare
- Where two haplogroups share a common ancestor, and when. mtdna/U5a2b2a3/compare
- Notable connections
- Historical figures on the same branch. mtdna/U5a2b2a3/notable
- Ancient connections
- Archaeological samples sharing your lineage. mtdna/U5a2b2a3/ancient
- Match time tree
- Your own matches placed on the dated tree. Requires a test and login; logged-out visitors see a demo. mtdna/U5a2b2a3/matches
Substitute any Y haplogroup and swap mtdna for y-dna to see the paternal equivalents.
For example: y-dna/A/classic
myOrigins 3.0 and the Chromosome Painter
Estimates global ancestry proportions and per-chromosome local ancestry jointly rather than in sequence, which keeps the segments and the percentages consistent and allowed the reference panel to grow from 24 populations to 90. I invented the method and authored the white paper.
How it works.
Family Finder Matching 5.0
Autosomal matching and relationship estimation that classifies whether a match pair is endogamous before predicting the relationship, which removes a systematic bias that inflates kinship estimates in founder populations. Co-authored.
How it works.
Archaic Origins
Neanderthal, Denisovan, and ghost-archaic ancestry as calibrated percentages and per-chromosome segments. Built and validated; release planned on Discover.
How it works.
Open source
Mitotree
The Mitotree release: tree files, haplogroup definitions, and the material needed to assign haplogroups against it. Distributed through HaploGrep v3 plugins, so an external lab can use it without changing its existing workflow.
github.com/genebygene/mitotree
Geo-genetic triangulation
Infers where someone’s ancestors lived by finding long DNA segments they share with people who have documented genealogies, then correcting for how heavily each region is represented in the reference data. Developed for the Beethoven genome study; the version used in that paper is archived at Dryad.
github.com/paulmaier/geo-genetic-triangulation
fasta2genotype.py
Converts RAD sequence data into any of several genotype formats, using whole-haplotype information rather than discarding all but one SNP per locus. Written for my dissertation, where standard pipelines were throwing away most of the signal, and recently updated for current Python.
github.com/paulmaier/fasta2genotype
Everything else is at github.com/paulmaier.
Tutorials
Complete, runnable walkthroughs in R, written for graduate students. Each one loads its own page.
Documentation and data
- myOrigins 3.0 white paper — full method description: phasing, segment classification, the conditional random field, phase correction, and how the 90-population reference panel was built.
- Family Finder Matching 5.0 white paper — the matching algorithm, endogamy classification, and relationship estimation.
- Anaxyrus hybrid panel — the diagnostic SNP panel and methods for identifying western toad and Yosemite toad hybrids.
Archived data and analysis scripts
- Yosemite toad phylogeography, gene pools, and connectivity — 4 GB covering three papers: phylogenetics, fastsimcoal2 demographic models, niche overlap, contact zones, hierarchical structure, AMOVA, hub and satellite analysis, isolation by distance, migration corridors, and ResistanceGA. Scripts included.
- Yosemite toad adaptive potential and future conservation units — outlier detection, gradient forest and generalized dissimilarity modeling, and the geo-genomic simulations behind the Geminate Evolutionary Unit framework.
- Yosemite toad transcriptome, speciation genes, and adaptive introgression — 3.4 GB covering the islands and rivers analysis across three replicate contact zones, the de novo larval transcriptome, and the tadpole growth-by-development GWAS.
- Geo-genetic triangulation — the script behind the biogeographic analysis in the Beethoven genome study.
- Raw sequence data — the reads behind all of the above.
PRJNA558546 — ddRADseq
PRJNA574353 — larval transcriptome
ON156774–ON156780 — L7 mitogenomes
PRJEB56343 — Beethoven genome
PRJNA510176 — Platycladus orientalis
Much of my production code is proprietary. Where that is the case I have linked the method description instead, which is detailed enough to reimplement. If you want to discuss any of it, or you are stuck reproducing something, get in touch.










