Skip to content
@jennifer/forensicgenetics

Limitations

Read this before quoting a number. Some of it will change what you do.

Not a validated forensic tool

The deck implements published estimators and is tested against them, but it has not been through the validation any jurisdiction requires of software used in casework. Treat it as a library for research, teaching and cross-checking, and validate independently before it informs a real conclusion.

Do not use this deck for a legal proceeding. It is not intended for, and is not recommended for, any calculation that informs a prosecution, a defence, a paternity ruling, an immigration decision, a disaster-victim identification, or any other legal or official determination. Jurisdictions require validated, accredited software for that purpose, operated under an accredited process; this deck is neither. If a number from here would ever reach a court, reproduce it with validated software and let that result be the one you rely on.

No warranty, no liability

This software is provided as is, without warranty of any kind, express or implied, including but not limited to the warranties of merchantability, fitness for a particular purpose, and non-infringement. You use it entirely at your own risk.

The author accepts no responsibility and no liability for any claim, damage, loss, or other consequence arising from this software or from any number it produces, however that number was reached and whoever relied on it. Verifying that a result is correct, appropriate, and lawful for your purpose is your responsibility alone.

This restates - and does not narrow, extend, or replace - the warranty and liability disclaimers of the LGPL-3.0-only licence this deck is released under. Where the two differ, the licence governs.

Everything below is a specific, technical limitation. None of it narrows the two statements above.

Data

The bundled frequency table is synthetic

testdata/frequencies.xml contains made-up numbers, generated only to exercise the reader and the examples. The file says so in its own header comment. Every figure printed by an example is therefore an illustration of the arithmetic, not a statement about any real population.

The examples build their table in examples/lib/exampledb.j; every number in it is made up. Download a real table from STRidER before computing anything you intend to rely on.

Pooling populations is not the same as an uncertain population

combineFrequencies merges tables into one. That merged table violates Hardy-Weinberg by construction whenever the inputs' frequencies differ - the Wahlund effect - so the match probability computed from it is biased downwards, which overstates the evidence.

For "the donor was German or Austrian", use randomMatchProbabilityAcross, which averages the match probability computed within each population rather than averaging the frequencies. On the worked example in the test suite the two differ by a third at a single locus, and the gap compounds across a profile.

Reserve combineFrequencies for the case where the merged group genuinely is the reference population.

Reading a published table is a separate deck

The core reads no published format. @mplx/strider converts STRidER-style XML, and carries its own limitations - chief among them that it has not been validated against a real STRidER export.

No reference sequence is bundled

lineage.align takes the rCRS (NC_012920.1) region you sequenced as an argument rather than shipping a copy. This also keeps the module database-independent, which is deliberate.

Statistical scope

Theta is not applied in kinship or pedigree calculations

kinship.j and pedigree.j use the plain product rule, θ = 0.

The sub-population correction does not extend to the IBD-1 term - the joint probability of two related people's genotypes under co-ancestry needs the full recursive sampling formula, not a swapped-in allele frequency. Applying θ to some terms and not others would produce a number that answers neither question, so those modules are uniformly theta-free.

θ is applied where it is exact: the single-source match probability. See Theta.

Kinship is for outbred pairs

The three IBD coefficients describe non-inbred individuals. An inbred pedigree needs the nine condensed Jacquard coefficients, which the pairwise form does not carry.

Peeling handles simple pedigrees only

A consanguineous mating, or any other loop in the individual/family graph, makes it cyclic. pedigree.validate rejects such a pedigree rather than mis-summing it. Loops need cutset conditioning, which is not implemented.

A half-sibling pedigree - one father, two mothers - is not a loop and is handled normally.

Mixture support stops at CPI / CPE

Mixtures gives the combined probability of inclusion and exclusion, which needs no assumption about contributor count but is correspondingly weak: it does not use the suspect's profile, ignores peak heights, and degrades badly under allele dropout.

Full deconvolution - multi-contributor models with drop-in, dropout and peak heights, the EuroForMix space - is a larger, later effort layered on this base. It is not here.

EMPOP's phylogenetic alignment rules are not applied

lineage.align renders a plain optimal alignment with indels slid 3'. EMPOP shifts indels to the position the mtDNA phylogeny prefers, which can differ for a complex indel. Check such a designation against EMPOP before quoting it.

The default mutation rates are placeholders

DEFAULT_MUTATION is rate = 0.002, decay = 0.1 - the right order of magnitude for paternal autosomal STR mutation, not a published figure.

This matters more than it looks. When a mutation model rescues a trio from a single inconsistency, the resulting CPI is directly proportional to the rate assumed at that locus, and observed per-locus rates vary over roughly an order of magnitude. Substitute published rates with withLocusRate.

Performance

Alignment cost grows with the region, and with divergence

align computes only a band around the diagonal, so for near-identical sequences - a haplotype against its own reference, which is the case it exists for - the cost is roughly linear in the region length rather than quadratic. Measured: 342 bp in 0.08 s, the ~1.1 kb control region in 0.29 s, 1600 bp in 0.39 s, all inside 50 MB. Time tracks length closely - 14.6x the region length for 15.7x the time, out to the 5000-base cap - which is the band doing its job.

Two sequences that genuinely diverge force the band wider, up to the full matrix, and then the cost is the quadratic one. align refuses a region longer than lineage.MAX_ALIGN_LENGTH (5000 bases) rather than appearing to hang; at that cap a near-identical pair takes about 1.2 s.

Everything else in the deck is fast. An eight-locus, four-typed-person pedigree LR runs in about a tenth of a second, and peeling memory is flat regardless of pedigree size.

Peeling slows as observed alleles accumulate

Cost is driven by the number of distinct alleles the typed members show at a locus, growing steeply in the resulting genotype count. Measured, for one locus: 4 alleles 11 ms, 8 alleles 90 ms, 12 alleles 0.6 s, 16 alleles 1.9 s, 20 alleles 5.4 s. A large reference family at a highly polymorphic locus can therefore take seconds per locus. Allele lumping already caps this at what is actually observed; beyond that, the panel is the thing to trim.

Building a database or profile one entry at a time is quadratic

The copy-returning builders take a copy of the whole container before writing into it, so calling one in a loop re-copies everything accumulated so far: withLocus over n loci costs 1 + 2 + ... + n entry copies. Measured, with the allele count held fixed at 10 so only the locus count varies: 25 loci in 5.9 ms, 100 in 78 ms, 400 in 1077 ms - each doubling costs about four times as much.

At the size a real STR panel has this does not matter. Fifteen to forty-five loci is tens of milliseconds, and converting a full 44-marker reference table lands around 16-21 ms.

profile(id, genotypes) takes the whole list and accumulates once, so build a profile that way rather than folding withGenotype over a list. There is no bulk counterpart for FrequencyDb yet: withLocus is the only public way to add a locus, and constructing FrequencyDb and LocusFrequencies directly to avoid it would skip the validation that catches a percentage-scaled table before it produces genotype probabilities above 1. That is not a trade worth making at this scale.

combineFrequencies used to pay this internally and no longer does.

Iterating a list of large values copies each element

for (def d in $dbs) binds a fresh copy of each element on every pass, and copying a FrequencyDb copies its whole locus map. Inside a loop over loci that is quadratic, which is what combineFrequencies was doing; indexing the list positionally reads the element in place and is linear. The same applies to any list whose elements carry a map or a long list - it is the container copy, not the loop, that costs.

Rotating two working buffers by assignment - spare = prev; prev = cur; cur = spare - copies for the same reason, once per rotation. Holding both in a two-element list and alternating an index avoids it entirely; that is what align does with its two score rows, and what keeps a banded alignment linear in the region length instead of quadratic.

The same trap catches lists.push and lists.pop, which return a fresh list rather than mutating: $xs[] = item appends in place, and a stack or queue is better walked with an index than rebuilt on every pass. The two pedigree graph walks do it that way.

Building a large pedigree is quadratic

withFounder and withChild return copies, as every builder in this deck does, so assembling an n-member pedigree copies the member map n times. That is invisible for the tens of members a real case has and noticeable in the hundreds. The likelihood functions themselves are unaffected.

Long panels underflow

randomMatchProbability and pedigreeLikelihood return ordinary products and will underflow on a very long panel. Use log10RandomMatchProbability, log10PedigreeLikelihood, or pedigreeLR, which takes the ratio locus by locus and stays in range.

Reporting

Numbers from this deck are conditional on assumptions the deck cannot state for you. Whatever you write down, write down alongside it:

  • which reference population, and where the table came from
  • which θ
  • which loci contributed - especially if a profile was narrowed
  • what the alternative hypothesis actually is (unrelated? a brother? an unknown man?)
  • for a probability of paternity, which prior
  • for a lineage marker, which database and how large it was