# TerrainLogic Scientific Validation Framework

**Public framework | August 13, 2026**

## 1. Claims are separated

TerrainLogic does not use one broad word—“validated”—for unlike evidence. Every claim is assigned to one of these levels:

1. **Implementation verification:** units, numerical methods, conservation, boundaries, and deterministic examples.
2. **Reference-model parity:** agreement with an applicable authoritative model or published reference calculation.
3. **Independent predictive validation:** performance against observations not used to choose methods, coefficients, defaults, or thresholds.
4. **Operational validation:** reliability, reproducibility, failure handling, state continuity, and report integrity in realistic workflows.

Passing a calculation check does not prove field accuracy. Matching a reference model proves compatibility with that model, not that either estimate equals a field observation.

## 2. Reproducibility and anti-leakage rules

- External sources are registered with citation, license, retrieval information, and integrity metadata.
- Source files remain immutable; normalized analytical tables preserve provenance.
- Dataset eligibility is decided before predictions are scored.
- Entire sites, studies, experiments, or years are held out where shared structure would make random-row splitting optimistic.
- Validation targets are never supplied as missing model inputs.
- Calibration evidence is not relabeled as independent validation.
- Missing required drivers cause exclusion or diagnostic classification—not target-derived repair.
- Passes, failures, exclusions, uncertainty, and applicability limits are published together.

Exact internal artifacts, adapters, test implementations, production architecture, correction history, and failed experimental formulations are confidential ContourStack R&D records.

## 3. Required statistics

Eligible numerical endpoints report sample size, mean bias error (predicted minus observed), PBIAS, MAE, RMSE, median absolute error, Pearson correlation, R-squared where appropriate, Spearman correlation, NSE, KGE where supported, log-scale sensitivity for skewed outcomes, and uncertainty coverage where observations provide uncertainty.

Correlation measures association, not agreement with the one-to-one line. It is never used alone as a pass criterion.

## 4. Whole-platform evidence matrix

| Capability | Current strongest evidence | Claim boundary |
|---|---|---|
| Sheet and rill erosion | National and locked RUSLE2 parity benchmarks plus deterministic verification | Planning agreement with RUSLE2; independent observed field validation remains open |
| Nutrient accounting | Mass conservation, independent crop and selected loss endpoints, treatment-direction tests | Planning-level screening and scenario comparison; not fertilizer recommendation or universal fate validation |
| Soil water | Water-balance and state-continuity verification | Modeled planning indicator; station/domain observational validation remains open |
| Soil Conditioning Index | Reference-calculation and sensitivity verification | RUSLE2-style planning index; not measured soil-carbon prediction |
| Gully hotspots | Terrain-processing and workflow verification | Inspection-priority screening; not confirmed gully occurrence or erosion rate |
| Grassed waterways | Independent numerical matrix and workflow verification | Preliminary planning only; hydrology, survey, outlet, FOTG, and professional approval required |
| Terraces | Independent numerical matrix and workflow verification | Preliminary planning only; alignment, hydrology, outlet, survey, FOTG, and professional approval required |

## 5. Current results

### Erosion

The national parity benchmark contains 749 eligible cases. Six hundred cases (80.1%) are within 20% of RUSLE2. Another 107 are outside 20% but within 0.25 ton/ac/year, producing 707/749 (94.4%) within the practical planning tolerance. Forty-two exceed both limits. Mean bias error is +0.475 ton/ac/year; raw annual A has Pearson r 0.9993 and R-squared 0.9986.

The separate locked 199-case regression suite has 176 cases (88.4%) within 20%; 195/199 (98.0%) fall within 20% or 0.25 ton/ac/year. Four exceed both limits. Mean bias error is -0.433 ton/ac/year; raw annual A has Pearson r 0.9973 and R-squared 0.9946.

These are reference-model parity results, not independent observed-erosion validation.

### Nutrients

The nutrient ledger passes mass-conservation and continuity checks. Independent evidence supports several crop-removal domains, a selected soil-derived total-phosphorus endpoint, one input-complete tile-nitrate diagnostic, ammonia placement direction, and management-response direction. Other endpoints—including generalized leaching, total denitrification, and universal residue phosphorus—remain failed, incomplete, or unsupported. The public nutrient validation report contains the endpoint scores and applicability limits.

### Preliminary engineering

The combined waterway and terrace verification matrix contains 50 independent calculation and invariant checks; all 50 pass. The evidence verifies calculation consistency within the tested numerical domain. It does not validate site hydrology, terrain alignment, outlet adequacy, utilities, state FOTG compliance, constructability, or a final field design.

## 6. Release and change-control rules

- A production change must pass the applicable focused tests and locked regression gates.
- Improvements to one case or endpoint must be checked for regressions elsewhere.
- Material changes require a new evidence fingerprint and updated public result date.
- Public claims cannot exceed the strongest completed evidence tier.
- Known limitations remain visible until new independent evidence resolves them.
- Detailed debugging, counterfactual methods, parameter decisions, and unsuccessful experiments remain in the confidential engineering record.

## 7. Current defensible conclusion

TerrainLogic has substantial implementation verification, strong RUSLE2 parity evidence for sheet and rill erosion, planning-level nutrient-accounting support with explicit process limitations, and verified preliminary engineering calculations. Independent field validation remains incomplete for several important endpoints. Public language must continue to distinguish verified calculations, reference-model parity, diagnostic evidence, and independent observational validation.
