A field-genetics case file · Morus rubra & Morus alba · eastern North America

The native tree that is
disappearing into its invader

And what it actually takes to detect that

Red mulberry is not mainly being cut down or crowded out. It is being absorbed — hybridised, generation by generation, into the introduced white mulberry that grows alongside it. In four studied populations in southern Ontario, more than half of the trees were already hybrids.1 Many of them look exactly like red mulberry.

So the practical question is not whether this is happening. It is how you would know — for a particular tree, a particular stand, or a particular batch of seedlings. That turns out to be a precise methods problem with a surprisingly rigid structure: what any method can detect is fixed by how the marker is inherited and by how many generations deep the hybridisation goes. Get that wrong and you can spend real money on a test that could never answer your question.

This page works through it in order. What the eye can do and where it stops. What sets the ceiling on any molecular method. An audit of the markers that already exist — and why, as of 2026, none of them transfers to this problem. Two new markers derived here from 45 published chloroplast genomes and 180 published gene sequences, one maternal and one biparental, both readable on a strip of agarose. Then how to choose a route, and how many trees you actually have to test.

53.7%trees that were hybrids
0usable published panels
2fixed differences derived here
1discordant reference plastome
computed here published unverified

Every substantive claim below carries one of these. Computed here means it was calculated directly from public sequence data while writing this page, and can be reproduced from the accessions given. Unverified means I could not confirm it against a source and you should not rely on it.

File 01 · The disappearance

A tree can go extinct without anything dying

Red mulberry (Morus rubra) is native to eastern North America, from Ontario and Vermont south to Florida and west to Texas and South Dakota. White mulberry (Morus alba) was brought from Asia for silkworm culture and is now one of the commonest weedy trees on the continent.

The two species hybridise readily, and the hybrids are fertile. Where they grow together, pollen moves overwhelmingly in one direction, because there is overwhelmingly more white mulberry pollen. Each generation of backcrossing dilutes the native genome a little further, and the endpoint is not a dead tree — it is a population of trees that still look more or less like red mulberry and are no longer genetically red mulberry.

The measurement everyone cites comes from Burgess, Morgan, Deverno and Husband, published in Molecular Ecology in 2005.1 They genotyped 184 trees from four populations in southern Ontario where both species grow together, using nuclear markers alongside chloroplast sequence.

What they found published
MeasureValue
Trees that were nuclear hybrids53% (98 of 184)
Pure red mulberry29% (53)
Pure white mulberry18% (33)
Range of hybrid frequency across the four sites43% – 67%
Hybrids with more white than red mulberry markers67%
Hybrids carrying a white mulberry chloroplast68% of 25

43 polymorphic RAPD fragments — of which only nine were species-diagnostic, five for white mulberry and four for red — plus chloroplast sequence from an 802 bp window of rbcL, in which the two species differ at just three fixed sites.1 Sampling was stratified: every putative red mulberry was taken, along with a roughly 25% subsample of the putative white and hybrid trees within 25 m of each. The authors note this may overestimate hybrid frequency and underestimate white.1 The paper's abstract gives the hybrid count as 53.7% (n = 99) while its results section and Figure 3 give 53% (n = 98); 53 + 33 + 98 = 184, so the figures here follow the results section. The widely quoted "53.7%" comes from the abstract.

That last line is the one that shapes everything that follows, and it is worth stating the other way round: 32% of those hybrids carried a red mulberry chloroplast. Hold onto it. It is the reason no chloroplast test — including the good one below — can ever be the whole answer.

How firm is that 32%? Less than it looks

The chloroplast result rests on 25 hybrids, not 184 — sequencing was done on 42 trees in total. Seventeen of the 25 carried the white mulberry chloroplast type. Burgess and colleagues tested that split against 1:1 and could not reject it (χ² = 2.72, P = 0.099), which is why their own discussion says most hybrids carried the white chloroplast type "although insignificantly".1

Run the interval and the honest range is wide: the fraction of hybrids with a red mulberry mother — the ones a chloroplast test cannot see — is somewhere between 17% and 52% at 95% confidence. computed here The point estimate is a third. The upper end is half.

This page uses "about a third" as shorthand throughout, because it is the best available estimate. Read it as a third, and possibly half. Nothing downstream should depend on the difference — and where something would, the page says so.

Michigan lists red mulberry as state threatened.2 In Ontario, where the species sits at its northern range edge and hybridisation pressure is worst, most known red mulberry sites have white mulberry growing in them.2

Why this needs a method at all

A red mulberry that is genetically half white mulberry will still make fruit, still feed birds, and still be counted as a native tree in a plant survey. The loss is invisible without a test, which means every number anyone quotes about how much red mulberry is left depends entirely on how the trees were called. Nobody is systematically screening the trees in any given county, and the trees are not going to be screened by looking at them.

Everything after this is about the second question — given that you need a test, which one, and what will it actually settle? The answer is more constrained than it first appears — worth understanding before you spend anything.

File 02 · The suspects

What you can actually see from the ground

The two pure species are genuinely distinguishable by eye. The trouble starts with everything in between — so it matters which characters actually separate the species and which are folklore.

CharacterRed mulberryWhite mulberryWorth?
Hairs on the leaf undersideErect hairs spread evenly over the whole blade — soft to the touchConfined to the main veins and the tufts in vein axilsBest character
Leaf areaBlade 10–18 cm, often much largerBlade 8–10 cmBest measurable
Upper leaf surfaceRoughened, dull greenGlossy, lustrousGood
Marginal teethSmall, numerous, pointedFewer, larger, bluntGood, underused
Leaf apexDrawn out to a long pointAcute to bluntSuggestive
Bark textureFlat thin plates peeling outwardsFirm braided ridges, orange showing in the furrowsSuggestive
Petiole lengthOverlapping; if anything M. alba is longerUseless
Style lengthBoth species effectively lack a styleUseless
Male vs female treesBoth subdioecious; ~10% hermaphrodite, and individuals switch between yearsUseless
Fruit colourBoth range from white through red to near-blackUseless

Compiled from the Flora of North America treatment,8 Nepal, Mayfield & Ferguson 2012,9 and Nepal 2008.10 published

The character to learn is where the hairs sit on the underside — not whether hairs are present. Red mulberry carries erect hairs spread across the whole blade, soft to the touch. White mulberry has them only along the ribs and in the little tufts where veins meet the midrib.

Three things widely believed that are not true

Fruit colour means nothing. Nepal and colleagues put it bluntly: fruit colour is "highly variable within M. alba and non-diagnostic. In fact, in wild populations, fruits of M. alba are usually red to black rather than white."9 The Flora of North America gives white mulberry syncarps as "black, purple, or nearly white."8 More confident misidentification traces to this one piece of folklore than to anything else.

Style length is not diagnostic. This one circulates widely in identification guides. Both species effectively lack a style — Nepal's genus-wide key places M. alba and M. rubra together in the short-or-absent-style half of the genus.10 What older sources call "style length" is a measurement of the stigma arms, and even that is contested.

Whether a tree is male, female or both tells you nothing. Both species are subdioecious, with roughly one tree in ten bearing both sexes, and individuals changing their expression between years.9 Treating breeding system as a species character is precisely the error that produced a spurious mulberry species, M. murrayana, later dismantled by exactly that observation.9

A disagreement in the sources, left open

Bark colour is not settled. The Michigan abstract calls red mulberry bark "dark-reddish brown",2 while the Flora of North America, the Canadian status report and Nepal all describe it as grey to greyish-tan and put the orange tint on white mulberry, showing in the furrows between firm ridges and on exposed roots.8129 The reproducible part is the texture — flat peeling plates against firm braided ridges — so use that and ignore the colour. unresolved

One more caution that undercuts almost everything in the table: these characters are read from mature leaves on ordinary shoots. Juvenile growth, stump sprouts and vigorous water shoots converge between the species, and as Nepal puts it, "nearly all of the unique characteristics of M. rubra fail in juvenile leaves."9

File 03 · Why the witness lies

Morphology sees species. It cannot see ancestry

Burgess and colleagues measured six morphological characters alongside their genetic markers, and found that the pure and hybrid classes differed on all six.1 That sounds like good news for field identification. It is close to the opposite.

The reason is in which classes the characters separate. Of the six, only leaf area and leaf perimeter told all three groups apart. For the other four — number of lobes, sinus depth, and trichome density on both leaf surfaces — white mulberry and the hybrids were statistically indistinguishable from each other, and both differed from red mulberry.1

Hybrids do not look intermediate. They look like white mulberry. In the canonical discriminant analysis, "M. alba and hybrid mulberry were more similar to each other than either was to M. rubra."1

This is better news than it sounds for one job and much worse for another. Morphology is a decent tool for finding the trees that are not red mulberry — which is what a removal programme needs. It is close to useless for the question of whether a particular good-looking tree is pure, because the hybrids that most resemble red mulberry are exactly the ones the characters fail on.

Trees that fooled the experts

The sharpest demonstration comes from a Kansas population studied by Nepal, where trees were first assigned by an expert on leaf, bud and bark characters and then genotyped.10 The morphological calls did not survive:

Marker systemTrees called pure by morphology that were genetically admixed
Microsatellites10 — nine of them called red mulberry, one white
RAPD markers9 — five called red mulberry, four white

Of nine trees that morphology had flagged as possible hybrids, only six were confirmed. And the two marker systems agreed with each other on only 44% of the hybrids they found — a reminder that even the genetic answer depends on which markers you use.10

Conservation practice has already absorbed this. Canada's recovery strategy designates critical habitat for trees "confirmed as pure-strain Red Mulberry trees through genetic testing," and lists confirming the genetic purity of morphologically-identified trees as outstanding work.13 The status report records the consequence plainly: "a few of trees previously counted as Red Mulberry were determined to be hybrids and were excluded from subsequent surveys."12

Morphology is a screening tool, not a verdict. It will correctly sort most pure trees. It cannot tell you that the tree in front of you is pure, which is the question actually being asked.

The one quantitative consolation: of everything measurable on a leaf, hair density on the underside is the best single predictor of what the genome actually says. Regressed on its own against hybrid index it explains about 30% of the variation, and in a multiple regression across all six characters it was the only one that stayed significant, with the whole model reaching 40%.1 Thirty percent is a good morphological character. It is not a test.

So morphology hands the problem to the molecules. The question is which molecules, and that is not a matter of taste.

File 04 · The ceiling

Two things decide what a test can possibly see

Before comparing methods on price or convenience, it is worth knowing that most of the answer is already fixed by two properties of the situation: how the marker is inherited, and how many generations of backcrossing have happened. Neither is negotiable, and together they rule out whole categories of test before you spend a dollar.

One: a maternal marker can only ever name one parent

Chloroplasts are inherited maternally in the great majority of flowering plants — the chloroplasts in a tree came from the ovule, not the pollen. Burgess and colleagues relied on exactly this when they used mulberry chloroplast DNA to establish which parent was the mother in each hybrid.1 It makes chloroplast DNA a superb species marker and a fundamentally limited hybrid marker.

If a red mulberry flower is pollinated by white mulberry, the seedling is a 50/50 nuclear hybrid carrying a pure red mulberry chloroplast. Every chloroplast test ever devised will call that tree red mulberry — correctly, and uselessly. This is not a hypothetical failure mode. It is about a third of the hybrids Burgess found — with a 95% interval running from 17% to 52%, because that result rests on 25 sequenced trees and its 68:32 split could not be distinguished from an even one.1 computed here No amount of money spent on a better chloroplast assay recovers those trees, because the information is not in the molecule.

What a chloroplast test is excellent at is the other direction. A tree that looks like red mulberry but returns a white mulberry chloroplast is definitively not pure, and you have found that out for the price of one PCR — polymerase chain reaction, a routine lab method that makes millions of copies of a target stretch of DNA, enough to work with. Here, that stretch carries the differences between red and white mulberry. Where white mulberry is the usual pollen donor and therefore often the maternal parent too, that catches most of the hybrids — just not all.

A consequence that changes how you sample

Because it is maternal and does not recombine, a chloroplast type is a property of a maternal lineage, not of an individual. Every seedling of one mother tree carries her chloroplast, so testing her second seedling tells you exactly nothing you did not learn from the first.

That has a sharp practical edge. If you are looking at a set of related trees — a seed lot, a batch of nursery stock, a stand of root suckers — the chloroplast assay counts mothers, not stems. One test per maternal family is the entire available information, and the highest-value move is not a better assay but keeping track of which seed came from which tree. Where lineage is unknown and material has been mixed, the same assay becomes informative again in a different way: it estimates what fraction of the mix has white mulberry mothers.

Two: detection decays by half with every backcross

Closing the maternal blind spot needs a marker inherited from both parents. But a biparental marker has its own ceiling, and it is arithmetic rather than chemistry.

At a locus where the two species are fixed for different variants, a first-generation hybrid carries one copy of each — heterozygous, unmistakable, detectable with certainty at a single locus. Backcross that hybrid to white mulberry and each offspring has a one-in-two chance of inheriting the red variant at that locus. Backcross again and it is one in four. With n independent fixed-difference markers, the chance a first backcross slips through looking pure is 0.5n:

Chance an admixed tree is scored as pure computed here
Fixed-difference markersFirst-generation hybridFirst backcrossSecond backcrossThird backcross
10%50%75%88%
100%0.1%5.6%26%
200%0.0001%0.3%6.9%
500%negligible0.00006%0.13%

The first column is the one people miss. In principle, a single biparental fixed difference detects a first-generation hybrid every time — no probability is involved, because an F1 inherits one copy from each parent and must therefore carry both variants. What one locus cannot do is see deep backcrosses. Ten markers catch first-generation hybrids and first backcrosses. Fifty make it unlikely that anything within three generations slips past. Thousands — which is what sequencing gives you — let you estimate the actual ancestry fraction rather than answering yes or no.

What "in principle" is carrying

That table describes an idealised single-copy locus: two alleles per individual, fixed between the species, both amplifying equally, both visible in the readout. Real markers fail those assumptions in specific ways, and each failure moves a tree from the left of the table towards the right.

It matters here because the nuclear marker this page ends up recommending is ITS, which is not single-copy. It is a tandem array of hundreds to thousands of repeats, and what a PCR returns is a pooled and potentially biased sample of them. An F1 is expected to show both parental repeat classes, but a class can be under-represented through copy-number differences between the parents, primer mismatch, competition during amplification, or partial homogenisation of the array. So the guarantee in the first column is a property of the genetics, not a measured property of the assay — and it has not been measured for this one. Where the page says a marker "detects every F1", read it as detects every F1 whose minority repeat class amplifies above the detection threshold.

The published guidance agrees. A simulation study by Vähä and Primmer concluded that efficient detection of first-generation hybrids needs 12 to 24 markers, but that "separating backcrosses from purebred parental individuals requires a considerable genotyping effort (at least 48 loci), even when divergence between parental populations is high."6

And the wild population sits at the wrong end of that table

It would be convenient if most hybrids in the field were F1s, because that is the column every method handles well. They are not. Burgess's hybrids had a mean hybrid index of 0.46 — measurably below the 0.5 an F1 would give — and 67% of them carried more white mulberry genome than red.1 The authors' reading is explicit: "some of the hybrids are not F1 crosses; rather they are later generation backcrosses that contain high proportions of M. alba genome."

That makes sense given the history — white mulberry arrived in the early 1600s and mulberry generations are short, under about 15 years, so there has been time for many rounds of backcrossing. But it means the class a single locus catches perfectly is not the class most trees fall into. A one-locus screen is strongest against the hybrids that are rarest in the field and weakest against the ones that are commonest.

This does not make such a screen worthless — a first-generation cross is exactly what you get from a verified pure mother in a white mulberry pollen cloud, so the F1 case is the common one in a controlled setting. But for a wild tree of unknown pedigree, expect the population you are screening to be mostly backcrosses, and price the answer accordingly.

Put the two together and the shape of the problem is fixed. A maternal marker rejects trees cheaply and can never clear one. A biparental marker reaches F1s and about half of first backcrosses, subject to actually detecting what is there. Certainty about deep ancestry needs tens to thousands of loci — which means sequencing, not a gel.

That is the yardstick for everything below. Any method is worth exactly what it resolves on those two axes.

File 05 · The audit

What already exists, and why none of it transfers

The obvious move is to find the published marker panel for this species pair and use it. I went looking, and what I found was several partial precedents and no finished one. That needs stating precisely, because "no panel exists" is easy to say and easy to get wrong.

The bar being applied

A marker set is ancestry-informative for this problem if it has been validated against three things at once: geographically representative M. rubra reference material, equivalent M. alba reference material, and known hybrids of known generation — F1s, reciprocal F1s, and backcrosses. Without the third, you cannot measure the two numbers that matter: how often the panel calls a real hybrid pure, and how often it calls a pure tree admixed.

Everything below is measured against that bar. Several of these efforts are good work; none of them clears it.

The foundational study used markers nobody can reuse

Burgess and colleagues' 2005 paper is still the reference measurement for mulberry hybridisation, and it is where the 53.7% comes from. The nuclear markers were RAPDs — randomly amplified polymorphic DNA — read alongside chloroplast sequence.1 RAPDs are dominant, so a heterozygote is indistinguishable from a homozygote for the present allele, and they are anonymous: the bands are not tied to known sequence. They are also notoriously sensitive to reaction conditions, which is why results generally do not transfer between laboratories. As a published measurement the study stands. As a protocol to pick up and run, it is not available.

The numbers underneath are worth knowing, because they set the resolution of the best measurement anyone has. They screened 100 RAPD primers and kept five. Those yielded nine truly species-diagnostic fragments — five for white mulberry, four for red — and 34 further fragments that were polymorphic but not diagnostic, for the 43 that were scored.1 Against the table in File 04, nine diagnostic markers is enough to catch F1s and most first backcrosses, and thin for anything deeper. Their reference material was not perfectly separated either: the red mulberry reference set scored 0.89 on a hybrid index where 1.0 is pure, and the white scored 0.09 rather than 0.1

The microsatellite work exists, and was not designed for this question

There are two relevant efforts. Nepal's 2008 dissertation ran both microsatellites and RAPDs on a Kansas population — the analysis behind File 03, where ten morphologically pure trees came back admixed.10 The more recent one is Schreier and Nepal's survey of 78 trees across six Upper Midwest populations, posted to bioRxiv in July 2026 and not yet peer reviewed.20 preprint They screened 12 markers originally developed in Asian MorusM. indica, which their paper treats as synonymous with M. alba although Kew currently accepts it as distinct, and M. boninensis. Five amplified cleanly in M. rubra and were used for the analysis.

A correction to an earlier version of this page

This section previously argued that because those markers amplify in both species, they cannot distinguish them. That is wrong, and it is worth saying so plainly. Microsatellites are scored by allele size, not by whether they amplify: to compare two species at a locus, the primers generally have to work in both. Diagnostic power comes from how the allele-size distributions differ, how strongly the parental populations are differentiated, and how many loci you combine. Cross-species transferability is what makes a locus testable in both species, not what disqualifies it.

The real limitation is narrower and the authors state it themselves. The study was designed as an M. rubra population survey, and its own limitations section names "the absence of reference M. alba and confirmed hybrid genotypes," concluding that "additional highly informative nuclear markers are therefore needed to resolve the extent, directionality, and demographic consequences of introgression."20 So whether those five loci are ancestry-informative is not answered in the negative — it is unmeasured, because the design could not measure it.

Where individuals showed extra allelic peaks — the pattern that looks most like introgression — the authors decline to call it: such profiles "do not independently demonstrate allopolyploidy or introgression from M. alba," and what is needed is "species-diagnostic nuclear SNPs or genome-scale data."20

One risk worth flagging without overstating it. Observed heterozygosity came in below expected at every locus in every population, a mean of 0.34 against 0.65, and the authors list null alleles and allele dropout among the possible causes.20 If a primer silently fails on one species' allele, a heterozygous hybrid can read as a homozygous pure tree. But a heterozygote deficit on its own does not establish that — inbreeding, population subdivision (the Wahlund effect) and sampling structure all produce the same signature, and the authors name those too. Since the dataset contained no confirmed hybrids, nothing here shows dropout actually converting hybrids into pure calls. It is a specific failure mode to test for with known crosses, not an observed outcome.

The older result is the more pointed one. In Nepal's data the microsatellite and RAPD systems agreed on only 44% of the hybrids they found.10 Two marker sets, one population, and the ancestry calls largely disagreed — which is the clearest available demonstration that an unvalidated panel produces answers whose reliability you cannot assess.

Two operational efforts, neither publicly specified

Canada runs the most serious red mulberry identification programme anywhere. Leaf samples go to the University of Guelph Arboretum for genetic analysis to determine whether each tree is red, white or hybrid, and the recovery strategy designates critical habitat only for trees "confirmed as pure-strain Red Mulberry trees through genetic testing."13 The programme has enough confidence in its calls to exclude trees from surveys on the strength of them.12 But neither the recovery strategy nor the Arboretum's published material names the markers, the loci or the technique. not disclosed in any source I could find

Separately, a US SARE-funded citizen-science project — Red Mulberry Search and Rescue — collected over 100 leaf samples nationally, had DNA extracted at the Ohio University Genomics Facility, and worked with Los Alamos National Laboratory to begin building an M. rubra genome to compare samples against the existing M. alba one.21 It reports classifying samples as rubra, alba or hybrid "with a high degree of accuracy" — while stating the limitation itself: because no complete rubra genome existed, "the results are not necessarily indicative of complete purity of species."21 That is a genome-comparison approach rather than a marker panel, and its classifier is not published in a form anyone can rerun.

Neither of these is a criticism. A conservation programme has no obligation to publish a protocol, and a farmer-grant project has no obligation to release a classifier. It does mean the two most operationally exercised methods for this exact question cannot be picked up by a third party.

The honest version of the finding: as of August 2026 I could not locate a published, portable, multilocus panel validated for ancestry classification against geographically representative M. rubra and M. alba references together with known F1 and backcross hybrids. Markers that might contribute to one exist — microsatellite, ITS and organellar (chloroplast and mitochondrial). Their sensitivity and specificity for hybrid detection remain to be established.

Which sets the target for what follows. The most useful thing to look for is a fixed difference — a position where every red mulberry carries one variant and every white mulberry another, tied to known sequence so anyone can check the claim, and readable without specialist equipment. The rest of this page derives two, one maternal and one biparental. Neither has been validated against known crosses either, and the page says so where it matters.

File 06 · Building a marker

Where the two genomes actually differ

Deriving a fixed difference used to mean a sequencing project. It no longer does — enough Morus sequence is already public that the marker can be found by measurement rather than by bench work, and checked by anyone who wants to repeat it.

In 2025 a group at South Dakota State University published complete chloroplast genomes for 45 mulberry trees collected across eight US states, deposited as GenBank accessions PQ309062–PQ309106.3 That dataset makes it possible to stop guessing and measure which piece of DNA to look at.

I downloaded all 45 and analysed them directly. computed here

Aligning a red mulberry genome (159,423 bp) against a white mulberry one (159,293 bp) gives 421 single-base differences and 696 separate insertion or deletion events. The largest single indel is 36 bp, and inspecting the ten largest shows most of them to be tandem-repeat expansions — a short motif repeated one extra time. Those expand and contract on their own, are prone to assembly error, and make unreliable species markers.

That rules out the laziest possible test. There is no big clean length difference you could see by running a PCR product straight onto a gel. The dependable signal is in substitutions, and substitutions have to be either sequenced or cut with an enzyme.

Which region carries the signal

For each candidate region I counted positions where all 33 unambiguous red mulberries were fixed for one base and all 10 unambiguous white mulberries fixed for another.

RegionFixed differencesof which substitutionsVerdict
rpl32–trnL(UAG)13114The one to use
ycf19324Strong but unwieldy
ndhFrpl326822Strong
psbEpetL5114Good
trnStrnG3611Usable
psbAtrnH102Too weak
trnLtrnF82Too weak
rbcL (standard barcode)55Works, barely
matK (standard barcode)44Works, barely

computed here from GenBank PQ309062–PQ309106.

Two results stand out. First, the standard plant barcodes do workrbcL and matK carry five and four fixed differences respectively. That is unusual for two species in the same genus, and it means a conventional barcoding workflow is not useless here. But with only four or five informative positions, one sequencing error costs you a quarter of your evidence.

That is also, in effect, what Burgess used. Their chloroplast work sequenced an 802 bp window of rbcL and found the two species differing at three fixed sites, with no variation within either species.1 Three sites across 42 trees was enough to call maternal lineage, and it is consistent with the five this analysis finds across the whole gene. The point of what follows is not that rbcL fails — it is that rpl32–trnL carries roughly forty times more signal, which is what makes a restriction digest possible instead of a sequencing run.

Second, rpl32–trnL(UAG) is in a different league at 131 fixed differences. It separated all 43 unambiguous trees perfectly: every red mulberry scored 131 out of 131 red-type positions, every white mulberry 131 out of 131 white-type. No intermediates, no ambiguity.

File 07 · Two false leads

The marker that wasn't, and the reference that misleads

A perfect 69 bp marker, which does not exist

Early in the analysis a different region looked ideal. The spacer between rps15 and ycf1 came out at 333 bp in every white mulberry and 405–407 bp in every red mulberry — a 70 bp gap, trivially readable on a gel, no enzyme needed.

It is an artifact. The ycf1 gene is annotated as starting 69 bp further along in the white mulberry records than in the red mulberry ones. The same physical DNA therefore falls inside the gene in one set of records and inside the spacer in the other, and comparing "spacer lengths" compares two different things. Searching all 199 Morus chloroplast genomes in GenBank for the supposedly red-mulberry-specific 69 bp block found it present in every single one, including all 101 white mulberries. computed here

Recorded here because it is an easy and completely invisible way to invent a marker. Any length difference derived from annotation coordinates rather than from the sequence itself deserves this check.

The NCBI reference plastome for red mulberry is a white mulberry plastome

NC_070233, the designated NCBI RefSeq chloroplast genome for Morus rubra, carries a white mulberry–type plastome, and should not be used as a representative red mulberry chloroplast reference.

Scored against the diagnostic positions it comes out 12 white-type to 2 red-type. Its length, 159,289 bp, sits with white mulberry (159,293 bp) and nowhere near the red mulberry range of 159,396–159,423 bp. GenBank records it as identical to accession OP161259.4 The authors of the 2025 study independently noticed that this accession falls among the Asian species in their phylogeny.3 computed here

Note carefully what that does and does not say about the tree it came from. Everything in File 04 applies here too: a plastome reports maternal lineage, not nuclear ancestry. The source specimen could genuinely be a largely M. rubra tree that carries a white mulberry chloroplast through introgression — the gene flow that repeated backcrossing produces — which is not a rare event, it is the majority of the hybrids Burgess sequenced. So the defensible statement is that the record is taxonomically discordant, not that the plant was misidentified.

Either way the practical consequence is the same. Anyone comparing a sample against "the M. rubra reference genome" is comparing it against a white mulberry–type plastome and will get the wrong answer. Use the vouchered PQ309073–PQ309106 series instead.

The same check turned up the mirror image. Accession PQ309072, deposited as M. alba, carries a 131-out-of-131 red mulberry–type plastome — a tree identified as white mulberry whose maternal line was red. computed here That is exactly the bidirectional introgression Burgess described, caught in a modern dataset by accident, and it is the cleanest illustration on this page of why a plastome names a mother rather than a species.

The general lesson

GenBank holds 249 nucleotide records for M. rubra against 5,097 for M. alba. computed here Some of the red mulberry records carry white mulberry–type sequence and some of the white carry red, which is what you should expect from two species that hybridise freely — the labels record what a collector determined, and an organellar sequence records something narrower. If you are going to compare your tree against a reference, check the reference first.

File 08 · The two assays

One PCR, one enzyme, one gel

The 131 fixed differences in rpl32–trnL(UAG) can be read by sequencing. They can also be read for a few dollars with a restriction enzyme, because some of those differences create or destroy an enzyme's recognition site.

I searched the amplicon — the stretch of DNA the PCR copies — for enzymes whose cut count differs consistently between the species, then computed the predicted fragments for all 43 unambiguous trees. One is close to ideal.

HpyCH4III

Every red mulberry carries one cut site in the amplicon. Every white mulberry carries three. computed here The resulting patterns are not subtle size shifts needing careful measurement — they are different pictures.

Predicted digest · 1.5% agarose 15001000800 500300200100 LADDER UNCUT RED MULBERRY WHITE MULBERRY ~1800 bp 912 + 900 770 + 731 + 191 + 187
Fragment sizes computed from 43 published chloroplast genomes; band positions plotted on a logarithmic migration scale. The two red mulberry fragments differ by 12 bp and will run as one heavy band. The diagnostic feature is the small white mulberry fragment near 190 bp, which red mulberry never produces.

Read it as a shape rather than a measurement. Red mulberry gives one heavy band high on the gel. White mulberry gives a band slightly below it plus an obvious small band near the bottom. A hybrid with a white mulberry mother gives the white mulberry pattern; a hybrid with a red mulberry mother gives the red mulberry pattern. The test reports the mother, and only the mother.

If HinfI is what you can get

HinfI is cheaper and more widely stocked, and also separates the two, though the pattern is busier. computed here

Fragments above 60 bp
Red mulberry679, 562, 228, 168, 126
White mulberry688, 591, 425, 126

The diagnostic difference is that red mulberry has bands near 228 and 168 bp where white mulberry has a single band near 425 bp. SspI and BfaI also work if those are what you have.

The second assay, and this one is biparental

Everything above reads the chloroplast, which puts it squarely against the first ceiling from File 04: it names the mother and nothing else, and is blind to the third of hybrids that had a red mulberry mother. Closing that gap needs a locus inherited from both parents — and there is one that can be run in the same afternoon, on the same DNA extraction.

ITS — the internal transcribed spacer of the ribosomal RNA genes — sits in the nuclear genome, so a tree inherits it from both parents. If the two species carry different ITS variants, a hybrid carries both at once, and both are visible on a gel. That is the codominant signal — both parental variants visible at once — that the arithmetic in File 04 requires: a first-generation hybrid must show it, so a single locus detects that class with certainty.

I pulled every full-length Morus ITS sequence from GenBank — 90 labelled M. rubra and 90 labelled M. alba — and aligned them. computed here There are 22 near-fixed differences between the species, and one of them creates a restriction site:

At one position, all 88 clean red mulberry sequences lack an MboI site and all 90 white mulberry sequences have one. A perfectly fixed nuclear difference.
Predicted ITS digest · MboI · 1.5% agarose 15001000800 500300200100 LADDER RED MULBERRY HYBRID WHITE MULBERRY 691691 + 491 + 186491 + 186
The hybrid lane is the point. Because ITS is inherited from both parents, a tree carrying both variants shows the red mulberry band and both white mulberry bands together — a result neither pure species can produce. Fragment sizes computed from GenBank ITS records; a 22 bp fragment common to all three is too small to see and is omitted.

The amplicon uses the universal ITS1 and ITS4 primers published by White and colleagues in 1990,16 which are the most widely used primers in plant molecular biology and match Morus directly. verified against the sequences The enzyme is NEB MboI, R0147S, 500 units for $88.00. confirmed

ITS1  TCCGTAGGTGAACCTGCGG
ITS4  TCCTCCGCTTATTGATATGC

Both parental variants have been recovered from a real hybrid

The prediction above would be worth little on its own. A group at the University of Central Missouri did a relevant experiment in 2010, and the result is sitting in GenBank — though it is important to be exact about what it does and does not show.

They took a herbarium-vouchered M. alba × M. rubra hybrid — specimen KANU:361918 — cloned its ITS, and sequenced four individual clones from that one tree.17 I downloaded all four and scored them at my diagnostic site. computed here

Clone from the one hybrid treeMboI sitesReads asSubmitters' own annotation
HQ144170 · clone 12white mulberry type"Morus alba haplotype"
HQ144171 · clone 22white mulberry type"Morus alba haplotype"
HQ144175 · clone 31red mulberry type"Morus rubra haplotype"
HQ144187 · clone 41red mulberry type"Morus rubra haplotype"
One tree, both parental ITS types, split cleanly by the marker — and labelled the same way by the original researchers, who reached that conclusion independently and fifteen years earlier.

For completeness I scored their pure reference trees too: all eight red mulberry clones carry one MboI site, all three white mulberry clones carry two. computed here The separation is total.

What this does not show

Those four sequences came from cloning: the ITS product was split into individual molecules, and each was sequenced separately. Nobody has taken bulk PCR product from a hybrid, cut it with MboI, and shown that all three predicted bands are visible on a gel.

The distinction is not pedantry. Cloning recovers a repeat class that is present at any abundance; a digest only shows one that is present at enough abundance to make a visible band. A hybrid whose red-type repeats have been partly outcompeted during amplification could sequence as mixed and still run as a clean white mulberry pattern. So this result establishes the biology — a real hybrid carries both parental ITS variants, at this exact site — and leaves the assay's sensitivity unmeasured.

A second variant at the same locus, from a published study

While checking this, I found a second polymorphism at the same locus. A 2025 paper in Plants genotyped 542 mulberry accessions across the ribosomal region, recovering 158 SNPs and 15 indels, and built a CAPS marker — a PCR-and-enzyme test like the ones here — on a 13 bp insertion in ITS1.18 In my own downloaded set that insertion is present in red mulberry and absent in white. The insertion creates an MstI site; FspI is an isoschizomer recognising the same TGCGCA sequence, and is the more commonly stocked of the two, so it is the one worth ordering.

Be clear about what that paper did and did not do, because it is easy to read it as more supportive than it is. Its CAPS assay was built to separate M. alba and M. notabilis from other Morus species — a taxonomic discrimination — using BstEII and MstI. It was not a red-versus-white hybrid test, it did not use FspI, and it did not measure whether the assay detects mixed parental repeat classes in an F1 or a backcross.18 What it does establish is that the locus carries real, scoreable variation and that a restriction assay on it works at the bench.

Scored against my own downloaded set, the insertion is present in 88 of 88 clean red mulberry sequences and 1 of 90 white mulberry. computed here So the ITS region carries two diagnostic variants, at different positions, each readable with its own enzyme.

Two variants, but not two independent tests

An earlier version of this page called these "two unlinked nuclear markers" and claimed that running both means a hybrid has to fail twice to be missed. That was wrong. The MboI site and the 13 bp indel sit in the same nuclear ribosomal repeat. They are physically linked, they travel together in the same tandem array, and they share nearly every way this assay can fail: concerted evolution, unequal repeat abundance between the parents, preferential amplification of one repeat class, or a minority class simply falling below the detection threshold.

Anything that hides one variant will tend to hide the other in the same tree. Running both gives you two observations of one locus, which is a useful check against a bad digest or a misread gel — but the errors are correlated, so it does not multiply your confidence the way two genuinely independent loci would. Independent confirmation has to come from somewhere else in the genome.

That paper also publishes mulberry-specific ITS primers, which are worth preferring over the universal ones if you are ordering fresh.

The caveat, now smaller than I expected

ITS sits in hundreds of tandem copies, and those copies can be homogenised over generations by concerted evolution — which would erase the hybrid signal. The honest worry was that this makes the test fail on older hybrids.

The evidence says homogenisation in Morus is incomplete. The 2025 survey of 542 accessions concluded that the "widespread occurrences of heterogeneous SNPs and InDels" indicate "incomplete concerted evolution of nrDNA" — cloning recovered 26 distinct ITS sequences from 32 clones of a single tree, and 15 to 26 unique sequences per plant in the others.18 A separate study found polymorphic ITS types in 14 of 33 accessions.19 The hybrid above had not been homogenised at all. No published work reports concerted evolution erasing the red/white distinction.

That is better news than I expected, but it is not a guarantee for a specific tree several generations into backcrossing. Treat a clean result as strong evidence against recent hybridisation rather than proof of purity.

One geographic limitation

Morus celtidifolia, the Texas or mountain mulberry of the southwestern US and Mexico, shares the red-mulberry-type insertion. computed here Within the eastern range of M. rubra the two do not meaningfully overlap, but in the Southwest this test cannot be assumed to separate them.

The same GenBank check turned up the now-familiar problem: two of the 90 sequences deposited as M. rubra carry pure white mulberry ITS at every diagnostic position. computed here They are OR251260 and FJ605516 — and the independent literature search reached the same two accessions by a different route.

The protocol

  1. Collect and dry

    Young, fully expanded sun leaves. Dry them immediately in silica gel at roughly ten times the tissue mass. This single step matters more than anything else in the workflow — properly dried tissue yields good DNA for years at room temperature, and a leaf left in a warm bag overnight may yield none. Photograph the tree, the bark and both leaf surfaces, and take a GPS point.

  2. Extract DNA

    A silica-column plant kit, or CTAB if you prefer to mix your own. Mulberry leaves are high in polysaccharides and phenolics, so add PVP to a CTAB prep or use a kit with an inhibitor-removal step. A generic animal-tissue kit will disappoint you.

  3. Amplify rpl32–trnL(UAG)

    Published universal primers from Shaw and colleagues.5 I checked both against the actual Morus sequences: computed here

    rpL32-F  CAGTTCCAAAAAAACGTACTTC
    One mismatch to Morus, which reads …CCG… where the primer has …CCA…. It sits at position 8, far from the 3′ end, and will amplify normally.

    trnL(UAG)  CTGCTTCCTAAGAGCAGCGT
    Exact match in both species.

    Expect roughly 1,800 bp. Run 5 µL on a gel to confirm a single clean product before digesting.

  4. Amplify ITS as well

    Same DNA, second tube, primers ITS1 and ITS4. Expect roughly 700 bp. Running both loci from one extraction costs one extra tube and doubles what the afternoon tells you.

  5. Digest

    Take 10 µL of each product, add buffer and about 5 units of enzyme — HpyCH4III for the chloroplast amplicon, MboI for the ITS amplicon — and hold at 37 °C for an hour. Both enzymes work at the same temperature, so they can share a water bath. No purification step is needed for a diagnostic digest.

  6. Run and read

    1.5% agarose, alongside a 100 bp ladder and — this matters on your first attempts — a known white mulberry as a positive control. White mulberry is everywhere; find one in a hedgerow and use it to prove your assay works before you trust it on anything rare.

    Read the two lanes together. The chloroplast lane names the mother's species. The ITS lane says whether both species are represented in the nuclear genome. Three bands in the ITS lane is the result you are looking for and hoping not to find.

Honest status of this assay

These are candidate screens with sequence-level support, not validated diagnostic tests. Keeping those apart is the whole point of this box.

What is solid. The underlying sequence differences. The chloroplast marker holds across 43 independently sequenced genomes with no exceptions; the ITS marker separates 178 sequences correctly, and both parental variants have been recovered from a genuine vouchered hybrid at the diagnostic site. The locus is real and the polymorphism is real.

What is not. Every fragment size on this page is predicted computationally, and neither digest has been run on a bench by me. unverified as a bench protocol More importantly, neither has been run against known crosses — F1s, reciprocal F1s, backcrosses — which is the only way to measure the two numbers that decide whether a screen is any good: how often it misses a real hybrid, and how often it flags a pure tree. Those numbers are currently unknown for both assays. The published CAPS work at the ITS locus18 shows that a restriction assay there works at the bench, but it was built for species-level taxonomy and never measured hybrid sensitivity.

So treat your first runs as validating the method rather than the trees — that is what the white mulberry control is for — and treat every confidence figure later on this page as conditional on a sensitivity nobody has measured yet.

File 09 · Doing it yourself

What the bench actually costs

Both assays can be bought as a service — see the next file — so owning the equipment is a choice rather than a requirement. It is worth making when you expect to run many samples, because the cost per tree collapses to a few dollars and the turnaround drops from weeks to an afternoon. Below what it costs, so the comparison is concrete.

Everything this needs was, twenty years ago, a university facility. It is now four appliances and a shoebox of reagents, and nothing below requires a licence, an institution, or an address that looks like a laboratory.

Prices marked confirmed were read off the vendor's own page on 2 August 2026. Everything else is an estimate and should be treated as such.

The four machines

  • Thermocycler $695 – $835

    The one genuinely non-negotiable instrument: it drives the PCR by cycling temperature precisely. miniPCR bio sells the mini8X at $695 and the mini16X at $835; they sell to individuals and the machines run off a laptop. confirmed15

    Used lab equipment is the cheap route — an older Bio-Rad or Eppendorf cycler on eBay or LabX typically goes for a fraction of that. They are heavy, loud, and completely adequate. price unverified

  • Gel electrophoresis rig with viewer $309

    Separates the cut fragments by size so you can read the pattern. The blueGel at $309 combines the tank, the power supply and a built-in blue-light transilluminator in one unit, which is what makes it worth the money over a bare tank. confirmed15

    Buying the cycler and the gel together as the miniPCR DNA Discovery System runs $950–$1,099 and is the single simplest purchase decision here. confirmed

  • Something that holds 37 °C $0 – $199

    For the enzyme digest. A dedicated incubator such as the Cozy Cube at $199 is tidy confirmed, but a kitchen sous-vide immersion circulator holds 37 °C perfectly well and most households that would attempt this already own one. You also want 65 °C for the DNA extraction, which the same device covers.

  • Micropipettes and tips $219

    You need to measure 1–20 µL accurately, and nothing in a kitchen does this. Three adjustable pipettes covering roughly 1–10, 20–200 and 100–1000 µL.

    The research-grade route is expensive — new Eppendorf, Gilson or Rainin three-packs run $1,130–$1,380. You do not need it. miniPCR's H-style adjustable pipettes are $59 each, so $177 for three, and Edvotek's are $95 each with a lifetime warranty. The tradeoff on the cheap ones is a three-month warranty, not accuracy you would notice here.

    Tips: miniPCR is the rare vendor selling genuine single 96-tip racks rather than boxes of 960 — $13, $13 and $16 for the three sizes, $42 the lot. Buying bulk from Bioland costs about $150 for 2,880 tips, which is far more than this project will ever use.

The reagents

  • HpyCH4III restriction enzyme $83

    The enzyme that does the actual discriminating. NEB catalogue R0618S, 250 units for $83.00, supplied with rCutSmart buffer, recognition site AC^NGT — which is exactly the site the genome analysis identified. At 5 units per digest that is 50 trees. confirmed14

  • MboI restriction enzyme $88

    The nuclear half of the test — the enzyme that reads ITS and can show a hybrid outright. NEB R0147S, 500 units for $88.00, site GATC, and it runs at 37 °C alongside the other digest. confirmed14 Sau3AI and DpnII cut the same site if that is what you can get.

  • HinfI restriction enzyme $77

    The alternative, if you would rather have an enzyme you will use for other things. NEB R0155S, 5,000 units for $77.00, site G^ANTC. Far more units for the money, busier band pattern. confirmed

  • PCR master mix $53

    NEB OneTaq 2X Master Mix, M0482S$53.00 for 100 reactions. A 2X master mix means you add only water, primers and template, which removes most of the ways a first PCR goes wrong. confirmed

  • 100 bp DNA ladder $71

    The size reference you read the gel against. NEB N3231S, 100 gel lanes, $71.00. confirmed

  • Four primers about $37

    Custom-synthesised oligos, ordered by typing in the sequence. Eurofins Genomics publishes $0.42 per base at the 25 nmol desalted scale — about $9.24 for a 22-mer, so roughly $37 for all four. Registration is an ordinary web signup. IDT sells to individuals too but shows no price without an account. One order lasts years. Two pairs: one for the chloroplast locus, one for ITS.

    rpL32-F   CAGTTCCAAAAAAACGTACTTC
    trnL(UAG)  CTGCTTCCTAAGAGCAGCGT
    ITS1       TCCGTAGGTGAACCTGCGG
    ITS4       TCCTCCGCTTATTGATATGC

    price unverified

  • Plant DNA extraction $22 – $359

    Mulberry leaves are loaded with polysaccharides and phenolics that inhibit PCR, so a generic animal-tissue kit will disappoint you. The column kits are the reliable option: Zymo Quick-DNA Plant/Seed, D6020, $273 for 50 preps; Qiagen DNeasy Plant Mini 69104 $326, or DNeasy Plant Pro 69204 $359. confirmed

    Much cheaper to start: miniPCR's X-Tract crude-lysate buffer is $22 for 20 extractions. A crude lysate is lower quality than a column prep, but for a robust multi-copy target like chloroplast or ribosomal DNA it is very often enough — and at roughly a dollar a tree it is the sane way to find out whether your protocol works before spending $273. Home-mixed CTAB with added PVP is the other cheap route.

  • Gel chemistry $19 – $85

    The shortcut worth knowing about: all-in-one agarose tablets that already contain the buffer and the stain$19 for 8 gels. Drop one in water, microwave, pour. For a first project that removes three separate purchases and the most tedious weighing step.

    Buying separately: miniPCR agarose 20 g $46 or GoldBio 100 g $132; TBE buffer powder $7.50 for 600 mL; Biotium GelRed 0.1 mL $34; 6X loading dye $30. confirmed

    Use GelRed, GelGreen, SYBR Safe or similar with a blue-light viewer — not ethidium bromide under UV. The modern stains are far less hazardous and blue light will damage neither your eyes nor your DNA.

  • 100 bp ladder $49 – $71

    Cheaper than the NEB one listed above if you shop: GoldBio ReadyLadder $49 for 500 µL, miniPCR's load-ready version $66. confirmed

  • Silica gel desiccant $79 for 55 lb

    The cheapest thing per unit on this list and the one that most determines whether any of the rest works. Buy the fine 0.5–1.5 mm non-indicating beads in bulk — a 55 lb drum is $79.20, about $3.18/kg confirmed — plus a small amount of indicating gel as a saturation cue. Avoid the 3–5 mm consumer beads sold for flower drying; drying speed depends on contact area with the leaf, and big beads have little.

What it comes to

A working bench, everything new: about $1,600.

LineChoiceCost
Thermocycler + gel rigminiPCR DNA Discovery System$950
Pipettes + tipsThree miniPCR H-style, one rack each$219
Both enzymesHpyCH4III + MboI$171
PCR master mixOneTaq, 100 reactions$53
Four primersEurofins, 25 nmol desalted$37
DNA extractionX-Tract buffer, 20 preps$22
GelsAll-in-one agarose tabs, 8$19
LadderGoldBio ReadyLadder$49
Silica gel55 lb drum$79
Total$1,599

Swap in a used thermocycler and gel rig and it lands nearer $800. Swap up to a proper column extraction kit and research-grade pipettes and it passes $2,500. The machines are the whole decision; everything else is noise.

Per-tree running cost after setup is a few dollars for both tests. The expensive part is the first tree; the hundredth is nearly free.

Before buying anything

Two cheaper routes are worth considering first.

Community biology labs already own all of this and will let members run their own projects. Verified monthly rates: BosLab, Somerville MA — $50; SoundBio, Seattle — $55, or $135 to lead your own project; ChiTownBio, Chicago — $75; Counter Culture Labs, Oakland — $100; BUGSS, Baltimore — $100; Genspace, Brooklyn — $110 community, $220 for full project access. confirmed Most ask for a short project proposal and a safety session.

Counter Culture Labs is the notable one for this project: it runs a standing plant biology group, and its fungal group already does ITS extraction and sequencing — which is precisely the second assay on this page.

Or send the reading out. Sequence the amplicon instead of digesting it, and you read all 131 chloroplast positions rather than one enzyme's worth. Psomagen charges $3.50 a reaction; Quintara from $4.00; Eurofins SimpleSeq prepaid kits work out at $6.10; Plasmidsaurus sequences an unpurified amplicon for $15 and takes a credit card with no minimum. confirmed Azenta/GENEWIZ confirms in its own FAQ that individuals can register and pay by card, though it publishes no price.

This does not remove the thermocycler

An earlier version of this page said sequencing lets you "skip the equipment entirely". It does not. Every service in that price band sequences a prepared template — a PCR product or a plasmid. None of them takes raw mulberry genomic DNA and returns the chloroplast region you asked about; the $15 Plasmidsaurus tier wants a linear amplicon, and SimpleSeq and Psomagen's standard Sanger service both assume you supply the product.

So the outsourced route still requires you to extract DNA, amplify the target on a thermocycler, check on a gel that you got one clean product, often clean it up, and supply a sequencing primer. For the 1,800 bp chloroplast amplicon you also need reads from both ends, since conventional Sanger runs out well short of that — so budget two reactions per tree, not one. The ~700 bp ITS amplicon fits in a single read.

What sequencing genuinely removes is the restriction digest and the gel readout: no enzymes, no ladder, no transilluminator, and far more information per tree. What it removes entirely is only true if someone else does your PCR — a community lab, a university core, or a collaborator.

So the fair comparison is not "sequencing versus the bench". It is the thermocycler plus a few dollars a read, against the thermocycler plus the gel rig plus enzymes. At three to fifteen dollars a read the sequencing route wins on information per tree and loses on turnaround, and it lowers the equipment bill by roughly the $309 gel rig and the $171 of enzymes rather than by the whole $1,599.

Sensible caution

None of this is dangerous work, but two habits matter: keep the DNA stain off your skin and out of the drain, and never run a sample you care about without a known control alongside it. The most common outcome of a first attempt is not a wrong answer — it is a blank gel, which tells you nothing and costs you a sample.

File 10 · Choosing a route

What each method actually buys you

Everything above can now be put on one axis. Each method costs something and resolves something, and the two are not proportional — the cheapest useful step is nearly free, and the last increment of certainty is the one that costs real money.

What each route resolves
MethodCost per treeDetectsBlind to
Leaf hairs, by eyefreeMost pure red mulberry vs everything elseHybrids, which resemble white mulberry
Chloroplast digest
rpl32–trnL, HpyCH4III
~$3White mulberry mothers — most hybrids, on a small sampleAny hybrid with a red mulberry mother — a third, perhaps half
ITS digest
MboI, biparental
~$3F1s, and about half of first backcrossesDeep backcrosses; any repeat class below the detection threshold
Sanger, both loci$7 – $60
plus PCR
Same, plus all 131 chloroplast positions and the exact ITS variantsSame generation limit — more sites, still one locus each
Genome skimming$225 – $1,000Ancestry fraction, hybrid index, generation classLittle — this is the answer if you need one

Per-tree costs assume you can already run a PCR; see File 09 for what a bench costs to build, and note that the Sanger row does not remove that requirement — those services sequence a prepared amplicon, so the thermocycler is needed either way. The chloroplast amplicon needs a read from each end at 1,800 bp, hence two reactions. Vendor prices verified below; the Sanger figures are unverified beyond the two checked by hand.

The useful reading of that table is that the two gel assays are not a cheap approximation of sequencing. They answer a different question. The chloroplast assay rejects trees; the ITS assay is aimed at exactly the class of hybrid a first-generation cross produces. What neither can catch is the tree several generations into backcrossing — red mulberry mother, ITS homogenised, and a quarter of its genome still foreign.

One column is missing from that table, and its absence matters more than anything in it: a sensitivity figure for each row. Nobody has measured how often these assays miss a hybrid that is genuinely there, because that requires known crosses and, per File 05, nobody has assembled that reference set for this species pair. The "detects" column describes what each method is designed to reach, not a measured hit rate. Read it that way, and read File 11 before attaching a confidence number to any of it.

For the deeply backcrossed tree, and for anyone who needs a number rather than a flag, you need the last row.

When you need certainty: genome skimming

Send extracted DNA for low-coverage whole-genome sequencing. At one or two times coverage you recover the complete chloroplast genome as a free byproduct — it is present in thousands of copies per cell — and enough nuclear variants to place the tree on a triangle plot of hybrid index against heterozygosity, which distinguishes first-generation hybrids from backcrosses and from pure trees.7 One experiment, both answers, no marker development.

What it costs, for a real tree, today: Plasmidsaurus publishes $250 for 1 Gb, $500 for 5 Gb, $1,000 for 15 Gb, and lists plants among the organisms the service covers. confirmed Read the tiers rather than the headline, though. The $250 tier is specified for genomes of 20–60 Mb; a 340 Mb mulberry sits in the 300–750 Mb band, which is the $1,000 tier. A gigabase against 340 Mb is about 3× coverage and would probably carry a hybrid index, but it is not the configuration they sell for a genome that size.

The obstacle is not the vendor. It is the DNA. Plasmidsaurus does not accept any intact tissues or organs from animals, plants, insects, fungi, etc. confirmed A leaf in an envelope is refused. Plant material is taken only as protoplasts with the cell walls stripped, or as high-molecular-weight genomic DNA you extracted yourself — RNase-treated, and never touched by phenol or chloroform.

SeqCenter is the same story approached from the other side. It posts nanopore ligation prices in the open — $150 for 300 Mb, $175 for 600 Mb, $225 for 1.2 Gb, library prep included confirmed — with no quote gate and no institutional account required, which makes it cheaper per gigabase than the tier above. But it wants 60 µL at 40 ng/µL of clean double-stranded DNA, and its own extraction service covers only certain microbes.

So the sequencing is genuinely cheap and buyable off a web page. Turning mulberry leaf into a tube of sequenceable DNA is the step that still needs a lab bench or a core facility willing to do the extraction for you, and that step — not the sequencer — is what stands between a private individual and a plant genome.

One complication: there is no red mulberry nuclear genome assembly. NCBI lists zero assemblies and six sequencing runs for the species, against four assemblies for white mulberry. computed here Nuclear reads therefore have to be mapped to the white mulberry reference, which is workable for ancestry estimation but introduces a mild bias and needs someone who knows what they are doing.

That gap is closing on two fronts. Schreier and Nepal state in their preprint that low-coverage genome-skimming data already exist for the species and that development of a high-quality nuclear reference genome is underway using PacBio long-read and Hi-C scaffolding data, alongside complete chloroplast and mitochondrial genomes from which hundreds of candidate organellar markers are undergoing validation.20 preprint Separately, the SARE citizen-science project has been working with Los Alamos National Laboratory towards an M. rubra assembly for the same purpose.21 Anyone starting this work now should check whether either has landed before mapping to white mulberry.

The honest summary: the gel assays will correctly reject most hybrids, and the ITS assay will catch recent ones outright. Neither can certify a tree as pure. Anyone selling certainty from a PCR is overselling it.

Which is a reasonable place to end up. A test that costs a few dollars and rules out most of the problem is worth having, and it is the difference between a shortlist and a guess. Send the survivors for sequencing.

File 11 · How many

Rejecting is cheap. Certifying is asymptotic

One question remains, and it is the one that decides what a screening programme costs: how many trees do you have to test? The answer depends entirely on which way you want to be wrong, and the two directions have wildly different price tags.

Suppose you are looking at a group that should be red mulberry — a stand, a seed lot, a batch of planting stock — and some unknown fraction p of it is admixed. Testing n individuals, the chance of catching at least one hybrid is 1 − (1 − p)n:

Chance of detecting at least one assay-detectable hybrid computed here
Testedp = 50%p = 30%p = 20%p = 10%p = 5%
387.5%65.7%48.8%27.1%14.3%
596.9%83.2%67.2%41.0%22.6%
1099.9%97.2%89.3%65.1%40.1%
20100%99.9%98.8%87.8%64.2%
30100%100%99.9%95.8%78.5%
60100%100%100%99.8%95.4%

The left-hand column is not a pessimistic scenario. It is roughly what Burgess measured — 98 of 184 trees, in populations that had been sought out because they held red mulberry.1 Canada's recovery strategy repeats that same 53.7% figure for its core populations,13 which almost certainly makes it the same measurement rather than an independent confirmation — but it does mean the number is what the recovery programme itself plans around. Where that is the true rate, five tests settle it 97 times in 100.

Two cautions on borrowing that number, both from the authors. The sampling took every putative red mulberry but only about a quarter of the surrounding white and hybrid trees, so they warn the figure "is an overestimate of the frequency of hybrids (and underestimate of whites)".1 Pulling the other way, it counts only hybrids that survived to be sampled — "hybridization rates for Morus may be even higher at the time of fertilization."1 So treat 50% as a plausible working figure for a stand where both species grow together, not as a measured constant, and certainly not as transferable to a stand you have not sampled.

Now the other direction. Suppose all n come back clean. What have you proved? Only an upper bound, and it comes down slowly:

If zero of n are flagged, the 95% upper bound on the detectable fraction computed here
Individuals tested510203060100
Upper bound45%26%14%9.5%4.9%3.0%
Both tables assume a perfect test. It is not one

Those figures are what you get if every hybrid among the trees you sampled is actually flagged. That is not this assay. Write the assay's sensitivity as s — the probability that a genuinely admixed tree comes back positive — and the detection formula becomes 1 − (1 − p·s)n. Everything in the first table shifts right, and everything in the second bounds only p·s, not p.

We already know s is well below 1, and the page has said why. The chloroplast assay misses every hybrid with a red mulberry mother — a third of them on the best estimate, and up to half at the edge of the confidence interval. The ITS assay misses deep backcrosses by simple arithmetic, and may miss some recent ones through repeat-class bias. Nobody has measured s for either assay, because that needs known crosses, which is exactly the validation File 05 says is missing across this whole field. Until somebody does, these tables bound the frequency of hybrids your test can see, and say nothing rigorous about the biological hybrid fraction.

A second assumption is buried in the exponent. Both formulas treat the n trees as independent draws. Siblings from one mother are not independent — they share her chloroplast entirely, and half her nuclear genome. Neither are root suckers, which may be one clone, or trees clustered in one thicket. Sampling twenty stems from one maternal family is nowhere near twenty independent observations, and the effective sample size can be a small fraction of the stem count. Spread sampling across mothers and across space, or discount n accordingly.

Five tests reject a badly contaminated group. A hundred clean tests still leave a 3% detectable-hybrid fraction on the table — and a larger unknown one below the assay's threshold. Detection is cheap; certification is asymptotic and never finishes.

So design the programme to reject, not to certify. Test a handful from anything you are suspicious of and act on the first positive. Save the deep, expensive methods for the small number of trees that survive screening and actually matter — the ones you intend to collect seed from, or protect, or propagate.

And say what you found rather than what you wish you could say. Not this stock is pure red mulberry, which is not provable by any method on this page, but something a result actually supports: mother tree chloroplast red-type, both ITS variants red-type, twenty offspring screened with no hybrids detected by an assay of unmeasured sensitivity — so the detectable hybrid fraction is under 14% at 95% confidence. That is a much weaker claim. It is also the one the evidence carries.

File 12 · Afterwards

What to do with a result

Voucher everything. A genetic result attached to a GPS point, a photograph set and a pressed specimen is evidence; the same result attached to a memory is an anecdote. Herbarium sheets can be deposited with a regional herbarium, and most are glad of material from a documented wild population.

Consider sending samples onward. Madhav Nepal's group at South Dakota State University generated the 45-genome dataset this page is built on, works directly on red mulberry hybridisation,3 and has recently posted a population-genetic survey of 78 red mulberries across Kansas, Iowa, Wisconsin and Nebraska as a preprint.20 Their sampling is concentrated in the central US; the eastern and southeastern parts of the range look thin. Contributing tissue from an uncovered population is a real contribution rather than a favour asked.

And if a genuinely pure population turns up, that is worth telling your state natural heritage program about, whether or not red mulberry is formally listed where you are.

Sources

References

  1. Burgess, K. S., Morgan, M., Deverno, L. & Husband, B. C. (2005). Asymmetrical introgression between two Morus species (M. alba, M. rubra) that differ in abundance. Molecular Ecology 14: 3471–3483. pubmed.ncbi.nlm.nih.gov/16156816
  2. Penskar, M. R. (2009). Special Plant Abstract for Morus rubra (red mulberry). Michigan Natural Features Inventory, Lansing, MI. mnfi.anr.msu.edu
  3. Adhikari, B., Parajuli, S. & Nepal, M. P. (2025). Reporting complete chloroplast genome of endangered red mulberry, useful for understanding hybridization and phylogenetic relationships. Scientific Reports. GenBank PQ309062–PQ309106. pmc.ncbi.nlm.nih.gov/articles/PMC12008415
  4. Zeng, Q., Chen, M., Wang, S., Xu, X., Li, T., Xiang, Z. & He, N. (2022). Comparative and phylogenetic analyses of the chloroplast genome. Frontiers in Plant Science 13: 1047592. Source of accession OP161259, from which RefSeq NC_070233 is derived. ncbi.nlm.nih.gov/nuccore/NC_070233
  5. Shaw, J., Lickey, E. B., Schilling, E. E. & Small, R. L. (2007). Comparison of whole chloroplast genome sequences to choose noncoding regions for phylogenetic studies in angiosperms: the tortoise and the hare III. American Journal of Botany 94: 275–288. Source of the rpL32-F and trnL(UAG) primers. doi.org/10.3732/ajb.94.3.275
  6. Vähä, J.-P. & Primmer, C. R. (2006). Efficiency of model-based Bayesian methods for detecting hybrid individuals under different hybridization scenarios and with different numbers of loci. Molecular Ecology 15: 63–72. pubmed.ncbi.nlm.nih.gov/16367830
  7. Wiens, B. J. & Colella, J. P. (2025). triangulaR: an R package for identifying AIMs and building triangle plots using SNP data from hybrid zones. Heredity. nature.com/articles/s41437-025-00760-2
  8. Wunderlin, R. P. (1997). Moraceae: Morus. In Flora of North America North of Mexico, vol. 3. Oxford University Press. efloras.org — genus key, M. rubra, M. alba
  9. Nepal, M. P., Mayfield, M. H. & Ferguson, C. J. (2012). Identification of eastern North American Morus: taxonomic status of M. murrayana. Phytoneuron 2012-26: 1–6. The clearest published character comparison, and the source of the statement that fruit colour is non-diagnostic. phytoneuron.net (PDF)
  10. Nepal, M. P. (2008). Systematics and reproductive biology of the genus Morus L. (Moraceae). PhD dissertation, Kansas State University. Source of the Konza Prairie morphology-versus-genotype comparison. krex.k-state.edu
  11. Burgess, K. S. (2004). The genetic and demographic consequences of hybridization in small plant populations. PhD thesis, University of Guelph. The fuller analysis behind Burgess et al. 2005. Listed for further reading rather than cited — on re-reading the paper, every claim this page had attributed to the thesis is present in the 2005 article itself and is now cited there. atrium.lib.uoguelph.ca
  12. COSEWIC (2014). Assessment and Status Report on the Red Mulberry Morus rubra in Canada. Listed under COSEWIC assessments on the species page. Species at Risk Public Registry — Red Mulberry
  13. Parks Canada Agency (2013). Recovery Strategy for the Red Mulberry (Morus rubra) in Canada. Species at Risk Act Recovery Strategy Series. Listed under Recovery strategies on the same species page. Species at Risk Public Registry — Red Mulberry
  14. New England Biolabs product catalogue, prices confirmed 2 August 2026: HpyCH4III R0618S, HinfI R0155S, OneTaq 2X Master Mix M0482S, 100 bp DNA Ladder N3231S
  15. miniPCR bio (Amplyus LLC) store, prices confirmed 2 August 2026. minipcr.com/store
  16. White, T. J., Bruns, T., Lee, S. & Taylor, J. (1990). Amplification and direct sequencing of fungal ribosomal RNA genes for phylogenetics. In PCR Protocols: A Guide to Methods and Applications, pp. 315–322. Academic Press. The source of the ITS1 and ITS4 primers, now used across plants as well as fungi. doi.org/10.1016/B978-0-12-372180-8.50042-1
  17. Nikaido, S., Salah, S. M., Ely, J. S., Wolf, H. J. & Raveill, J. A. (2010). Mulberry (Morus: Moraceae) hybridization in eastern North America: morphological and molecular evidence. University of Central Missouri. Unpublished; data deposited as GenBank HQ144170–HQ144187, including four cloned ITS sequences from the vouchered hybrid KANU:361918. ncbi.nlm.nih.gov/nuccore/HQ144170
  18. Xu, X., Zhang, L., Bi, C., Qin, M., Wang, S., Li, D., He, N. & Zeng, Q. (2025). Extensive nrDNA polymorphism in Morus L. and its application. Plants 14(16): 2570, published 18 August 2025. Open access. 158 SNPs and 15 indels across 542 accessions. Source of the 13 bp ITS1 indel and of the evidence that concerted evolution in Morus is incomplete. Its CAPS assay uses BstEII and MstI and was built to separate M. alba and M. notabilis from other Morus species — a taxonomic discrimination, not a red–white hybrid test. doi.org/10.3390/plants14162570
  19. Xuan, Y., Wu, Y., Li, P., Liu, R., Luo, Y., Yuan, J., Xiang, Z. & He, N. (2019). Molecular phylogeny of mulberries reconstructed from ITS and two cpDNA sequences. PeerJ 7: e8158. doi.org/10.7717/peerj.8158
  20. Schreier, S. J. & Nepal, M. P. (2026). Population genetics of native red mulberry at its northwestern boundary suggests postglacial founder effects. bioRxiv preprint, posted 14 July 2026, not certified by peer review. 78 M. rubra individuals across six populations in Kansas, Iowa, Wisconsin and Nebraska. doi.org/10.64898/2026.07.11.737963
  21. Cornett, J. (2022–2024). Red Mulberry Search and Rescue: preserving genetic diversity for the future of sustainable agroforestry. USDA SARE Farmer/Rancher grant FNC22-1338. Citizen-collected leaf samples, extraction at the Ohio University Genomics Facility, genome comparison against M. alba with Los Alamos National Laboratory. Reports rubra / alba / hybrid calls but states the results are "not necessarily indicative of complete purity of species." projects.sare.org/sare_project/fnc22-1338