| Transect | Species count (richness) | Species |
|---|---|---|
| 3 | 2 | Halodule, Thalassia |
| 6 | 2 | DR: Chondria, Thalassia |
| 7 | 5 | DR: Acanthophora, DR: Chondria, Halodule, Syringodium, Thalassia |
| 8 | 3 | DR: Acanthophora, Halodule, Thalassia |
| 9 | 4 | DR: Chondria, DR: Hypnea, Syringodium, Thalassia |
| 10 | 1 | Halodule |
How Report Card Scores Are Calculated
Overview
Seagrass transects are monitored in Tampa Bay each year by multiple resource management agencies. To support these efforts, the Tampa Bay Estuary Program coordinates an annual intercalibration training prior to the field season. This intercalibration effort ensures that all participating groups are using consistent methods and interpretations. Seagrasses are independently assessed by each group at the same locations and the results are compared using report cards. The intent of these report cards is to highlight areas where each group is doing well and where efforts could be improved to ensure consistency. This document describes how the report cards are prepared and what the scores mean.
Because there is no single external ground truth, scores are based on how consistently groups agree with each other. A group that reports values close to the cross-group average for the year earns a high score and one that deviates substantially earns a lower score.
Scores are calculated for three field measurements for each seagrass species:
- Abundance: Braun-Blanquet (BB) cover category (0 = No cover, 0.1 = Solitary, 0.5 = Few, 1 = <5%, 2 = 5-25%, 3 = 25-50%, 4 = 51-75%, 5 = 76-100%)
- Blade Length: average blade length in cm
- Short Shoot Density: shoots per m²
Whether each group correctly identified which species were present is also factored into the abundance score. These three metric scores are averaged into an overall total score, which is then converted to a letter grade.
Note that “transect” is used herein to describe the intercalibration sites during training, which are quadrats at fixed locations. For the actual transect sampling, sites are set meter marks along a transect line.
The scoring workflow, step by step
The full scoring workflow, from raw field data to a final letter grade, is:
- Identify the “true” (consensus) species at each transect.
- Calculate the “true” value for each metric at each transect, as the cross-group average.
- Calculate the cross-group variation in reported values for each species/transect/metric, for the transect-level weight.
- Calculate each group’s difference (reported minus true) for each species/transect/metric.
- Apply the transect-level weight from step 3 to each transect’s difference from step 4, discounting differences at transects where groups disagreed a lot.
- Combine the weighted per-transect differences from step 5 across transects into one difference value per species/metric, using a true weighted average (weighted by the same transect-level weight from step 3, not a plain average), so a transect with a larger weight counts for more.
- Roll up species into a raw metric score as the plain average of those per-species differences.
- Establish the typical spread of raw metric scores between groups within a year, using a fixed baseline period of 5 years (2020-2024), this is the baseline every year (including years in the baseline itself) is compared against to judge whether that year was unusually tight or loose. The baseline is fixed rather than recomputed as new years are added, so a report card’s calibration doesn’t shift after the fact.
- Calibrate the score floor for that year and metric based on how tight or loose the cohort was that year relative to the typical spread from step 8, then rescale the raw score to a 0-100 numeric score.
- Get the overall Total score as the plain average of the three metrics’ numeric scores, then convert every numeric score, per-metric and Total, to a letter grade using fixed thresholds.
The list above reads top to bottom, but the steps aren’t a single chain: step 3’s weight is reused twice (in step 5 and again in step 6), and step 8 is a separate calibration pipeline built once from the fixed baseline years, not something computed along the way for a given group, it only joins back in at step 9. The diagram below shows those connections.
The sections below work through each of these ten steps in detail. Section headings are grouped to match how the steps are connected, so a heading spanning several step numbers (e.g. “Steps 3-7”) covers all of them together rather than one at a time, ending with a full worked example applying all ten steps to one group.
Step 1: The Consensus Species List
Not every species recorded across all groups counts toward scoring. A species at a given transect (site/quadrat) is considered present if at least two distinct groups reported it with non-zero abundance. This removes likely misidentifications while still being inclusive, i.e., a species does not need to be unanimous, just corroborated. Species identifications include seagrass and macroalgae. Both are scored on Abundance, since correctly identifying a species as present or absent matters regardless of species type (see Species Identification Penalties below), but only seagrass is used to assess Blade Length and Short Shoot Density, which are not measured for macroalgae.
The consensus list is created per transect, so a species may be on the list at one transect but not another.
Step 2: True Values
For each consensus species at each transect, true values are calculated as the cross-group average:
- Abundance: group Braun-Blanquet values (0, 0.1, 0.5, 1, 2, 3, 4, 5) are averaged directly across groups, then snapped to the nearest valid BB value (no coverage, solitary, few, <5%, 5-25%, 26-50%, 51-75%, 76-100%).
- Blade Length and Short Shoot Density: simple means across groups.
Only consensus species enter this calculation. This prevents a misidentified species (recorded by only one group) from distorting the true values for everyone else.
| Transect | Species | Abundance (0, 0.5, 1, 2, 3, 4, 5) | Blade Length (cm) | Short Shoot Density (per m²) |
|---|---|---|---|---|
| 8 | DR: Acanthophora | 1 | NA | NA |
| 8 | Halodule | 3 | 21.3 | 5.5 |
| 8 | Thalassia | 2 | 28.5 | 1.9 |
Steps 3-7: Group Differences and Metric Scores
For each metric (Abundance, Blade Length, Short Shoot Density), a group’s reported values are compared to the true values species by species and transect by transect, then rolled up step by step into a single score per group per metric. Before any of that rollup, though, a report can go wrong in a different way: reporting a species that isn’t really there, or missing one that is. Those species identification errors are handled first, since they change what counts as a “difference” in the first place.
Species Identification Penalties
Comparing a group’s reported values to the true values across all consensus species and transects reveals two types of species identification errors:
Missed species: a species is present but the group did not record it. The group’s abundance for that species is treated as “no coverage” (BB value 0) when computing the difference. The penalty scales with the true abundance of the missed species: missing a rare species is a small error, while missing a dominant one is a large error.
False positives: the group recorded a species that is not on the consensus list (i.e., no other group confirmed it). The true abundance is treated as “no coverage” (BB value 0), but unlike a missed species, the penalty is a fixed one category difference, not scaled by how high the group’s reported abundance was. What category a group happens to report for a species nobody else saw is not evidence about whether the species was really there, so it should not make the penalty any larger or smaller. A false positive is still an error and still counts against the score, it just counts the same regardless of whether it was reported at “solitary” or “76-100%” cover.
These penalties only affect the abundance score. Blade Length and Short Shoot Density cannot be meaningfully penalised for species that were not found.
For Abundance, the difference is computed as a gap in ordinal category positions rather than in raw Braun-Blanquet numeric values. The eight BB categories are treated as equally spaced steps (1 = no coverage through 8 = 76–100%), so a difference of 1 means the group was one category apart from the true value, regardless of the numeric gap between those categories on the BB scale. This implicitly assumes that there is similar difficulty in distinguishing between lower abundances (e.g., solitary vs few) and higher abundances (e.g., 51-75% vs 76-100%).
| Transect | Species | Abundance reported | Abundance true | Blade Length reported | Blade Length true | Short Shoot Density reported | Short Shoot Density true |
|---|---|---|---|---|---|---|---|
| 8 | DR: Acanthophora | — | <5% | NA | NA | NA | NA |
| 8 | DR: Laurencia | few | — | NA | NA | NA | NA |
| 8 | Halodule | 5-25% | 25-50% | 15.2 | 21.3 | 4.7 | 5.5 |
| 8 | Thalassia | 5-25% | 5-25% | 21.8 | 28.5 | 1.3 | 1.9 |
Per-transect difference and weight
At each transect, a difference is discounted when the different groups had a lot of trouble agreeing with each other there. That disagreement reflects how hard the species was to pin down at that specific site. It is not the evaluated group’s own performance. This is the only place a weight is applied. An earlier version of this methodology also tried to weight species by how consistent a group’s own difference was for that species across transects, but with typically only 2-4 transects per species per group that estimate was mostly sample-size noise (see Combining species into a metric score below), so species are instead combined with a plain average.
At transect \(t\), the difference for species \(s\) is
\[ d_{s,t} = \text{reported} - \text{true} \]
for Abundance (an ordinal category-position difference, see Species Identification Penalties above) and Blade Length (a raw difference in cm), or the symmetric percent difference
\[ d_{s,t} = \frac{\text{reported} - \text{true}}{(\text{reported}+\text{true})/2} \]
for Short Shoot Density (see Absolute vs. percent difference below). The one exception is a false positive transect for Abundance (see Species Identification Penalties above), where \(d_{s,t}\) is fixed at 1 rather than computed from the reported category, since what category was reported is not evidence about whether the species was really there. Its weight is
\[ w_{s,t} = \frac{1}{1+\sigma_{s,t}} \]
where \(\sigma_{s,t}\) is the standard deviation, across all groups, of their individually reported values at that transect for that species:
\[ \sigma_{s,t} = \text{SD}_g\!\left(x_{g,s,t}\right) \]
with \(x_{g,s,t}\) denoting group \(g\)’s reported value for species \(s\) at transect \(t\), and the standard deviation taken across all reporting groups \(g\), not just the group being scored. It measures how much trouble groups had agreeing with each other right there. A transect where fewer than two groups reported a value has an undefined \(\sigma_{s,t}\) and gets full weight (\(w_{s,t}=1\)). This is independent of how much the true value itself might vary from one transect to another. That spatial variability plays no role anywhere in scoring.
Rolling up to a species value
The species-level difference is a true weighted average, across transects, of the per-transect differences:
\[ \bar{d}_s = \frac{\sum_{t=1}^{T} w_{s,t}\,d_{s,t}}{\sum_{t=1}^{T} w_{s,t}} \]
A transect’s influence on \(\bar{d}_s\) scales with its weight \(w_{s,t}\), so a transect discounted for poor cross-group agreement is proportionally down-weighted relative to the other transects, rather than simply contributing a smaller value while every transect still counts equally toward the denominator.
When weighting matters
The weight \(w_{s,t}\) makes a difference count for less when groups had a lot of trouble agreeing with each other at that transect. Whether that raises or lowers \(\bar{d}_s\) compared to a plain (unweighted) average of the same transects depends on whether the group’s large differences happened to fall on the hard transects (high cross-group disagreement) or the easy ones (low disagreement). The two scenarios below use one hypothetical species reported at two transects to illustrate both cases.
Scenario 1: a large difference on a hard transect is discounted, lowering the score
A group is close to the true value at an easy transect (all groups agreed closely there, \(\sigma_{s,t}=0\)) but far off at a hard transect (groups disagreed a lot, \(\sigma_{s,t}=3\)).
| σs,t | |ds,t| | ws,t / average | |
|---|---|---|---|
| Transect A (easy, σ = 0) | 0 | 0.5 | 1.00 |
| Transect B (hard, σ = 3) | 3 | 4.0 | 0.25 |
| Weighted average | 1.20 | ||
| Unweighted average | 2.25 |
Without weighting, the two transects average to 2.25. With weighting, Transect B’s large difference is discounted because groups naturally struggled to agree there, yielding a weighted average of 1.2, a substantially better score. The group missed a hard target, and the scoring system appropriately reduces the penalty for that miss.
Scenario 2: a large difference on an easy transect is not discounted, raising the score
Now the same group is far off at the easy transect but close at the hard one.
| σs,t | |ds,t| | ws,t / average | |
|---|---|---|---|
| Transect A (easy, σ = 0) | 0 | 3.0 | 1.00 |
| Transect B (hard, σ = 3) | 3 | 0.5 | 0.25 |
| Weighted average | 2.50 | ||
| Unweighted average | 1.75 |
Without weighting, the two transects average to 1.75. With weighting, the small difference at the hard transect counts for less, so it no longer offsets the large miss at the easy transect, yielding a weighted average of 2.5. This is the appropriate outcome, since missing badly on a transect where all groups should agree is a genuine measurement error and should not be diluted by being close on a transect that is naturally harder to pin down.
The weight \(w_{s,t}\) does not systematically produce better or worse scores than a simple average. Whether it is higher or lower depends on the data. What it always does is give more say to transects where cross-group agreement was good, and less relative say to transects where even the group of observers as a whole struggled to agree, regardless of whether the scored group’s own report happened to be close or far off there.
Combining species into a metric score
The metric score is a plain average, across species, of each species’ absolute difference:
\[ \text{metric score}_{\text{raw}} = \frac{1}{S}\sum_s |\bar{d}_s| \]
where \(S\) is the number of species with a defined \(\bar{d}_s\) that year. No further weighting is applied here. An earlier version of this step weighted a species by how consistent the group’s own difference was for it across transects, on the reasoning that a species a group measured erratically should count for more. In practice, with typically only 2-4 transects available per species per group, that consistency estimate could not be separated from ordinary sample-size noise, and in several real cases it was driven almost entirely by how much the true value happened to vary across the few transects sampled, not by anything about the group’s own performance. Rather than weight species by an unreliable number, every species now counts equally.
Absolute vs. percent difference
\(d_{s,t}\) is not defined the same way for every metric:
- Abundance uses the raw difference in ordinal Braun-Blanquet category positions (as described in Species Identification Penalties above). These categories are already unitless and applied the same way to every species, so there is no reason to convert them to a percent basis.
- Blade Length uses the raw difference in cm.
- Short Shoot Density instead uses the symmetric percent difference, averaged across transects (after the per-transect weight, and again after species are combined). This metric is measured in physical units (shoots/m²) where the typical magnitude varies a great deal by species (e.g., Halodule is characteristically much denser than Thalassia) and by how dense the selected quadrats happened to be that year. A raw difference of a few shoots/m² means something very different for a sparse bed than a dense one, and years with higher or lower density across the selected quadrats would otherwise look artificially more or less consistent across groups even when relative agreement hadn’t changed. The percent difference puts every species and every year on the same relative scale, so a training year is not judged as unusually “tight” or “loose” just because the selected locations were denser or sparser that year.
Blade Length is also measured in a physical unit (cm) that varies by species, but it is deliberately not scored on percent difference. That split is its own topic, covered in Why the difference basis differs by metric below.
The weight \(\sigma_{s,t}\) follows the same basis as \(d_{s,t}\) for each metric:
- Abundance and Blade Length use \(\sigma_{s,t}\) directly, in raw units (category positions, cm).
- Short Shoot Density instead uses \(\sigma_{s,t}\)’s coefficient of variation, dividing it by the true value at that transect, for the same reason \(d_{s,t}\) does.
When \(\sigma_{s,t}\) cannot be computed (fewer than two groups reporting at that transect), the weight defaults to its neutral value, full weight (\(w_{s,t}=1\)), for every metric.
Why the difference basis differs by metric
Percent (or CV) normalization is only the right choice when the typical size of a measurement disagreement itself scales with the magnitude of the true value, a proportional, multiplicative error. If disagreement is instead roughly constant regardless of magnitude, an additive error, dividing by the true value doesn’t remove a scale effect, it creates one: the same absolute disagreement looks smaller wherever the true value happens to be larger, for reasons that have nothing to do with how consistent groups actually were.
The figure below tests this directly using every species/year combination across the full training record. Each point is one species in one year: its true value (x-axis) against the mean absolute disagreement between groups that year (y-axis), both in the metric’s native units.
For Short Shoot Density, disagreement climbs steeply with density (r = 0.97): a bed with more shoots present gives groups more opportunity to disagree on the count, the same pattern seen in count data generally. Percent difference correctly cancels that out, so a year or species with a naturally denser bed isn’t penalized just for being dense.
The figure below confirms that the cancellation actually works: it repeats the same species/year points for Short Shoot Density, but with the y-axis switched to the percent (CV) difference actually used for scoring, instead of the raw shoots/m² difference shown above.
Short Shoot Density’s percent difference shows a far weaker relationship with density (r = -0.33) than the strong one in raw units (r = 0.97) above. The CV basis is doing most of what it’s meant to: a dense bed and a sparse bed now contribute much more comparably to the score than they would in raw units. The residual negative trend is concentrated at the lowest true densities, a handful of species/year points below about 3 shoots/m², where a small true value in the denominator inflates the percent difference even for a modest raw miscount. That is a different, harder-to-eliminate artifact from the one motivating the move away from raw units in the first place, and worth keeping in mind when interpreting Short Shoot Density scores in years with especially sparse beds.
For Blade Length, disagreement is essentially flat regardless of blade length (r = 0.09): however consistent, or not, groups are at measuring a 10 cm blade, they are about equally consistent measuring a 30 cm one. This is closer to a fixed measurement precision (e.g., where along a tapering or curled blade someone chooses to measure) than a process that scales with the plant. Dividing by the true value here would not remove a magnitude effect, it would manufacture one. That is exactly what produced 2020’s oddly tight-looking Blade Length scores under an earlier version of this methodology that scored Blade Length on percent difference: that year happened to have unusually long blades, which mechanically shrank the percent differences even though the raw cm disagreement was unremarkable.
Based on this, Blade Length is scored on the raw cm difference (abs) rather than percent difference, while Short Shoot Density keeps the percent/CV basis (pct). Abundance, already on a unitless ordinal scale, is unaffected either way.
| Species | Reported (avg) | True (avg) | Mean difference (d̅s) |
|---|---|---|---|
| DR: Acanthophora | solitary | <5% | -1.8 |
| DR: Chondria | no coverage | few | -1.7 |
| DR: Hypnea | no coverage | few | -2.0 |
| DR: Laurencia | <5% | no coverage | 1.0 |
| Halodule | 5-25% | 25-50% | -1.2 |
| Syringodium | <5% | <5% | -0.5 |
| Thalassia | <5% | 5-25% | -0.6 |
To see how these combine into a raw metric score, consider the Abundance values for group C in Table 6 above. Taking the absolute values of the mean differences, the 7 species have \(|\bar{d}_s|\) of 1.8, 1.7, 2, 1, 1.2, 0.5, 0.6. The raw Abundance score is their plain average:
\[\frac{1}{S}\sum_s|\bar{d}_s| = \frac{1.8 + 1.7 + 2 + 1 + 1.2 + 0.5 + 0.6}{7} = 1.257\]
matching the Abundance value shown in the “Raw score” row of Table 11 in the full worked example below. This raw score is then scaled to a numeric score on a 0–100 scale using the range of scores across groups, which is then converted to a letter grade (see next steps).
Steps 8-9: Score Calibration
Raw metric scores are converted to a 0–100 scale. Without calibration, the best group in any year always maps to 100 and the worst always maps to 50 (a fixed minimum score), regardless of how closely groups agreed. This means a year where everyone performed very well would still produce a spread from A to D, which needs to be accounted for to avoid an unfair outcome.
To address this, the score floor (the minimum possible score) is raised in years when all groups agree closely with each other, and kept at 50 in years when disagreement is high.
How calibration works
For each year and metric, we compute the within-year standard deviation of group differences to assess the spread between groups, using the same difference basis as the metric’s score (percent difference for Short Shoot Density, cm difference for Blade Length, category-position difference for Abundance). This is then expressed as a ratio to the mean spread over a fixed baseline period, the first 5 years of training data (2020-2024). A ratio below 1 means groups agreed more closely than the baseline and a ratio above 1 means more disagreement than the baseline.
The baseline period is fixed rather than recalculated from all years seen so far. If it were recalculated every year, adding a new season of data would shift the baseline and could retroactively change the calibrated score floor, and therefore the grade, on report cards already published for earlier years. Fixing the baseline keeps a given year’s score floor stable once it’s been set.
Scoring Short Shoot Density on percent difference matters here in particular. It is measured in a physical unit whose typical scale varies by species and quadrat locations over the multi-year training record. Computing this spread on raw units would confuse a change in that scale with a genuine change in how closely groups agree, making some years look artificially tighter or looser than they really were. Percent difference keeps the year-to-year comparison on a consistent, scale-free basis. Blade Length is scored on raw cm differences instead, exactly because it does not have this problem (see Why the difference basis differs by metric above): its typical disagreement does not scale with blade length, so a percent basis there would introduce year-to-year scale drift rather than remove it.
The score floor for a given year and metric is:
\[ \text{floor}_{\text{year}} = \max\!\left(50,\ 50 + \left(1 - \frac{\text{SD}_{\text{year}}}{\overline{\text{SD}}}\right) \times 50\right) \]
where \(\text{SD}_{\text{year}}\) is the within-year spread for that metric and year, \(\overline{\text{SD}}\) is the baseline mean described above, and 50 is a scaling constant that sets the maximum possible floor lift (grade-points). The scaling constant is a subjective choice. Larger values compress scores toward each other in tight years, while smaller values preserve more spread. The maximum lift occurs when all groups agree perfectly (ratio = 0), giving a floor of \(50 + 50 = 100\), which corresponds to a minimum grade of A (all groups performed exactly the same). A year at the baseline average (ratio = 1) receives no adjustment and keeps a floor of 50. Loose years (ratio > 1) are capped at 50 so that no extra penalty is applied beyond the standard range.
Before reducing each year down to a single spread number, it helps to see the raw distribution that number comes from. The figure below shows every group’s raw (pre-rescale) metric score, one point per group, for every training year.
The next figure reduces each of those distributions down to a single within-year standard deviation and compares it to the fixed baseline mean (dashed line).
Effect on scores: tight vs. loose year
The table below compares the calibrated score floor for each year and metric, illustrating how the floor shifts in tighter training years.
| Year | Abundance | Blade Length | Short Shoot Density |
|---|---|---|---|
| 2020 | 50 | 81 | 69 |
| 2021 | 69 | 50 | 58 |
| 2022 | 50 | 52 | 50 |
| 2023 | 60 | 50 | 50 |
| 2024 | 52 | 50 | 56 |
| 2025 | 60 | 55 | 75 |
| 2026 | 60 | 91 | 59 |
Step 10: Letter Grades
After calibration, each group’s numeric score for each metric falls on a 0–100 scale, with the lowest possible score in a given year defined by the calibrated score floor. These are mapped to letter grades using fixed thresholds:
| Grade | Score range |
|---|---|
| A | 95 – 100 |
| A- | 90 – 94 |
| B+ | 85 – 89 |
| B | 80 – 84 |
| B- | 75 – 79 |
| C+ | 70 – 74 |
| C | 65 – 69 |
| C- | 60 – 64 |
| D+ | 55 – 59 |
| D | below 55 |
The Total score is the unweighted average of the Abundance, Blade Length, and Short Shoot Density numeric scores, then converted to a letter grade using the same thresholds.
Worked Example
The example below applies the general workflow from the Overview (steps 1-10) to group C in 2025, starting from the raw per-transect data.
Raw differences
The table below shows the raw data behind everything that follows: each row is one transect (site) and seagrass species, with the group’s reported value, the cross-group true value, and the difference between them for each metric (a percent difference for Short Shoot Density, a cm difference for Blade Length, a category-position difference for Abundance). The Diff column is already discounted by the transect-level weight \(w_{s,t}\) from Steps 3-7 (a transect where groups had more trouble agreeing with each other counts for less). The “Average” row for each species is a true weighted average of the transect-level values above it (the sum of Diff \(\times\) \(w_{s,t}\) divided by the sum of \(w_{s,t}\)), the same quantity used in scoring (Steps 3-7), so it is not a plain average of the Diff column shown here, nor is it what you’d get by averaging the Reported and True columns down to a single number and then subtracting, which is why those two columns are left blank on the Average row. The illustration below works through both points, and the weight itself, for one species.
| Transect | Abundance | Blade Length | Short Shoot Density | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Reported | True | Weighted Diff | Reported | True | Weighted Diff | Reported | True | Weighted Diff | |
| DR: Acanthophora | |||||||||
| 7 | 4 | 4 | 0 | — | — | — | — | — | — |
| 8 | — | 4 | — | — | — | — | — | — | — |
| Average | — | — | -1.8 | — | — | — | — | — | — |
| DR: Chondria | |||||||||
| 6 | — | 2 | — | — | — | — | — | — | — |
| 7 | — | 3 | — | — | — | — | — | — | — |
| 9 | — | 3 | — | — | — | — | — | — | — |
| Average | — | — | -1.7 | — | — | — | — | — | — |
| DR: Hypnea | |||||||||
| 9 | — | 3 | — | — | — | — | — | — | — |
| Average | — | — | -2 | — | — | — | — | — | — |
| DR: Laurencia | |||||||||
| 8 | 3 | — | — | — | — | — | — | — | — |
| 9 | 4 | — | — | — | — | — | — | — | — |
| Average | — | — | 1 | — | — | — | — | — | — |
| Halodule | |||||||||
| 3 | 5 | 7 | -1.17 | 20.6 | 24.6 | -0.88 | 12 | 9 | 23% |
| 7 | 5 | 6 | -0.58 | 20.8 | 5.1 | 1.41 | — | — | — |
| 8 | 5 | 6 | -0.67 | 15.2 | 21.3 | -1.08 | 4.7 | 5.5 | -12% |
| 10 | 5 | 6 | -0.59 | 5.6 | 1.7 | 0.85 | — | — | — |
| Average | — | — | -1.2 | — | — | 0.4 | — | — | 7% |
| Syringodium | |||||||||
| 7 | 4 | 5 | -0.58 | 23.8 | 7 | 1.09 | — | — | — |
| 9 | 3 | 3 | 0 | 18.4 | 20.9 | -0.36 | 0.3 | 0.7 | -53% |
| Average | — | — | -0.5 | — | — | 3.5 | — | — | -80% |
| Thalassia | |||||||||
| 3 | 2 | 3 | -0.53 | 6.2 | 23 | -1.5 | 0.7 | 1 | -24% |
| 6 | 5 | 5 | 0 | 17.4 | 4.9 | 1.13 | 0 | 0.2 | -53% |
| 7 | 5 | 6 | -0.59 | 33.8 | 9.1 | 1.27 | — | — | — |
| 8 | 5 | 5 | 0 | 21.8 | 28.5 | -0.85 | 1.3 | 1.9 | -26% |
| 9 | 5 | 6 | -0.58 | 22.6 | 24.4 | -0.4 | 3.3 | 3.3 | 0% |
| Average | — | — | -0.6 | — | — | -0.6 | — | — | -40% |
| Diff is weighted by how much groups disagreed with each other at that transect for that species and metric (Steps 3-7), so it is not simply Reported minus True. The Average row is a true weighted average of Diff (sum of Diff × weight, divided by the sum of weight), not a plain average. See the worked illustration further below for that weight applied to one species. | |||||||||
For example, group C’s short shoot density for Thalassia was reported at 4 transects in 2025:
| Transect | Reported | True | Percent difference | Cross-group CV (σs,t) | Weight (ws,t) | Weighted % difference |
|---|---|---|---|---|---|---|
| 3 | 0.7 | 1.0 | -35% | 0.45 | 0.69 | -24% |
| 6 | 0.0 | 0.2 | -200% | 2.78 | 0.26 | -53% |
| 8 | 1.3 | 1.9 | -37% | 0.43 | 0.70 | -26% |
| 9 | 3.3 | 3.3 | 0% | 0.10 | 0.91 | 0% |
Averaging the raw “Percent difference” column above gives -68%. Two further steps separate that number from what actually feeds into this species’ score. First, computing a single percent difference from the averaged Reported and True columns instead (i.e. (mean(Reported) - mean(True)) / ((mean(Reported) + mean(True)) / 2)) gives -19% instead. That is a different number, because it lets transects with a larger true value dominate the average instead of letting each transect count equally. Steps 3-7 always compute the per-transect difference first, as shown here, never this way. Second, each transect’s percent difference is weighted by the Weight column, based on how much trouble groups had agreeing with each other at that transect (its cross-group CV). That weight is applied as a true weighted average (Steps 3-7), not a plain average of the “Weighted % difference” column: dividing the sum of that column (-1.03) by the sum of the Weight column (2.56) gives -40%, matching the Average row for Thalassia’s short shoot density in Table 9 above and the value that actually feeds into this species’ difference, before the further species-level weighting described in Steps 3-7.
Scores
For each metric, the mean absolute difference per species (\(|\bar{d}_s|\)) is listed in a single comma-separated cell. The “Raw score” row below it shows the plain average of those values (computed using the formula in Steps 3-7), the calibrated score floor, the rescaled numeric score, and the letter grade.
| Species |d̅s| | Raw score (average) |
Score floor | Numeric score | Letter grade |
|---|---|---|---|---|
| Abundance | ||||
DR: Acanthophora (1.8), DR: Chondria (1.7), DR: Hypnea (2), DR: Laurencia (1), Halodule (1.2), Syringodium (0.5), Thalassia (0.6) |
— | — | — | — |
Raw score (average) |
1.257 | 60 | 66.6 | C |
| Blade Length | ||||
Halodule (0.4), Syringodium (3.5), Thalassia (0.6) |
— | — | — | — |
Raw score (average) |
1.500 | 55 | 95.5 | A |
| Short Shoot Density | ||||
Halodule (0.07), Syringodium (0.8), Thalassia (0.4) |
— | — | — | — |
Raw score (average) |
0.425 | 75 | 77.6 | B- |
| Total | ||||
Average of three metric scores |
— | — | 79.9 | B- |
| The Raw score row is the plain average of |d̅s| across species. |d̅s| is a percent difference for Short Shoot Density, a cm difference for Blade Length, and a category-position difference for Abundance, already discounted at the transect level for cross-group disagreement (Steps 3-7). | ||||
The figure below shows how raw scores map to numeric scores across all groups for each metric. The grey line is the linear rescaling anchored at 100 for the best group and at the score floor for the worst. Group C is highlighted in blue.
How the calibration affected this group’s scores
| Metric | Score without calibration | Score with calibration | Ratio (within-year spread / baseline mean) | Score floor |
|---|---|---|---|---|
| Abundance | 57.5 | 66.6 | 0.79 | 60 |
| Blade Length | 95.0 | 95.5 | 0.90 | 55 |
| Short Shoot Density | 54.8 | 77.6 | 0.50 | 75 |
A ratio below 1 (tighter than average cohort) raises the floor for all groups, including this one. A floor of 50 means no calibration adjustment was applied.