2026.2 A deterministic synthetic corpus for the amorphous-carbon Raman inverse problem2026.4 A release gate for modelled results, and the provenance that backs it

Finding 2026.3Learned inversion of modelled carbon Raman spectra, with a refusal gate

R. J. YorkFounder & CEO, SSX360 Corp. — Honolulu, Hawaiʻi

Released 2026-09-03 · Updated 2026-09-05 · Version 1.1 · Claim status MODELLED · Keywords: Raman spectroscopy, inverse problem, amorphous carbon, crystallite size, gradient boosting, refusal, out-of-distribution, linear probes, benchmark

Abstract. On the held-out split of the synthetic corpus of Finding 2026.2, an ensemble of gradient-boosted trees recovers crystallite size to 2.85% mean relative error and assigns the amorphisation stage with macro-F1 0.986, where the physics baseline fails at the turnover (901% mean relative error); a saturation detector used as a refusal signal lowers the error on corrupted rows without flagging a single clean one. Every number is a modelled result and supports no claim about real films.

Errata

2026-09-05 — v1.1: Table 2 restated on corpus v2 with the repository's own code (baseline row 0.211 / 9.01 / 2.38, T5 89.4 Å; CNN row from the revision-3 release record, 0.977 / 0.082 / 0.075 / 0.030 / 12.70); cross-model agreement 98.3% (release record); the 205,147 despiking count restated as points replaced, not spike events; the attribution probe restated as one representative query; the linear probes restated as in-sample fits on fixed representations of the training rows. No pipeline definition changed.

All errata

Contents
§2026.3(i)Purpose
§2026.3(ii)Pipelines compared
§2026.3(iii)Results on the held-out split
§2026.3(iv)Robustness and refusal
§2026.3(v)Probes: what the models use
§2026.3(vi)Verification
§2026.3(vii)Limitations

§2026.3(i) Purpose

This finding reports what the current inversion pipeline recovers from the synthetic corpus of Finding 2026.2, so that the same numbers can be re-run, unchanged, when measured spectra replace the synthetic ones. The results are a regression baseline for the pipeline, not a measurement of any material. Their value lies in three places: the gap between a physics-only baseline and a learned inversion at the Tuinstra–Koenig turnover; the behaviour of the pipeline on corrupted rows, where refusal rather than silent correction is the design rule; and probes that show what the models actually rely on.

§2026.3(ii) Pipelines compared

The physics baseline inverts the Tuinstra–Koenig relation, equation 2026.2.1, and assumes stage 1 throughout; it is the honest statement of what the disorder ratio alone can do. Against it are compared a gradient-boosted tree model on the subsampled spectrum together with the process metadata and the excitation wavelength; a one-dimensional convolutional network on the spectrum with a wavelength channel and a metadata branch; the tree model after a twenty-trial hyperparameter search; and an ensemble of the tuned tree model and the network with the mixing weight fitted on the validation split. The ensemble is the pipeline's current release.

Preprocessing was made robust in the same revision: a moving-median despiking step replaced 205,147 points across the 3,400 spectra (the count is of points replaced, not of spike events, because the rule also fires on ordinary noise excursions in the brighter parts of clean traces), and a detector-saturation check was added whose output is a flag, never a correction. A flagged row is refused, not repaired.

§2026.3(iii) Results on the held-out split

Scores on the 600-sample test split. T2 is the mean relative error in , a dimensionless fraction (the baseline's 9.01 is a 901% error), with the degenerate subset near the turnover reported separately; T3 is the mean absolute error in sp³ fraction; T5 is the mean absolute error in , in ångström, on the corrupted rows.

PipelineT1 macro-F1T2 rel. errorT2 degenerateT3 sp³ MAET5 corrupted, Å
Physics baseline (Tuinstra–Koenig, stage 1 assumed)0.2119.012.3889.4
Gradient-boosted trees, spectra + metadata0.9920.0330.0280.0155.14
1-D convolutional network, spectra + metadata0.9770.0820.0750.03012.70
Gradient-boosted trees, tuned0.9860.0290.0230.0144.09
Ensemble of tuned trees and network (release)0.9860.02850.0230.0193.48

The tuned tree model, the network and the ensemble are from the release evaluation record (pipeline revision 3); the untuned tree model is from revision 2.1, before despiking and the saturation flag were added, on the same corpus and split. T3 is not defined for the baseline, which returns no sp³ estimate; its T5 of 89.4 Å is the error on the corrupted rows before any gate.

The baseline's relative error of 9.01 is not a misprint: applied to a stage-2 film it returns , so the overestimation factor is , a factor of 8 at 10 Å and 37 at 6 Å — the practical content of the turnover. The ensemble is selected on T2 and is slightly worse than the tuned tree model on T3; that trade is recorded rather than hidden. The tuned tree model's hyperparameters were chosen by a search whose candidates were fitted on the training and validation rows together and scored on the validation rows, so the selection score was optimistic; the selected setting was refitted on the training rows alone, and every number above is from models that never saw the test split.

§2026.3(iv) Robustness and refusal

With the saturation flag applied as a refusal gate, the ensemble's error on corrupted rows falls from 3.48 Å to 3.01 Å, and no clean row is refused. A coverage probe that scores each spectrum against the training distribution flags 1.5% of the test split as out of distribution; relative error is 0.030 in distribution and 0.097 out of it, 20% of corrupted rows are flagged and 0% of clean rows. Both signals cost nothing on good data, which is the property required of a gate that will later stand in front of measured spectra.

§2026.3(v) Probes: what the models use

Linear probes — ridge regressions (sp³ fraction, scored by ) and multinomial logistic regressions (stage, scored by macro-F1) fitted in-sample on the 2,400 training rows for four fixed representations — test whether the quantities of interest are linearly encoded, and where. With the representations standardised: peak-only features give 0.73 / 0.81, the binned spectrum 0.99 / 1.00, process metadata alone 0.78 / 0.85, and everything together 0.99 / 0.99; with the labels shuffled the probes fall to chance ( ≤ 0.01, macro-F1 0.28–0.33 against 0.33 for three balanced classes). A nearest-neighbour attribution on one representative test query (14.2 Å, stage 2, 325 nm) finds its eight nearest training rows to be same-stage neighbours at 12.9–14.7 Å, and refitting the tree model without those eight rows moves the prediction by 0.014 Å (14.02 to 14.01 Å); the probe is an illustration on one query, not a population statistic. No single cluster carries the model.

The reading is cautionary as well as reassuring. Stage and are encoded in the spectrum itself. The sp³ fraction, by contrast, is partly reached through process metadata and through the stress- and hydrogen-induced shifts of the G band, which means that a visible-excitation Raman spectrum alone should be treated as a suspect source of sp³ fraction. That is consistent with the literature on ultraviolet excitation for sp³ [FR01] and it is the kind of statement the measured phase will have to confirm or retire.

§2026.3(vi) Verification

The results were admitted through the release gate of Finding 2026.4: the corpus regenerated byte-identically; retrained metrics matched the ledgered evaluation exactly; stage assignments of the tree and convolutional models agreed on 98.3% of samples against a required 95%; each optimisation step improved or held the score it targeted; and the probe sanity checks held. The evaluation records are published with the corpus at the public data record, and an independent regeneration on a second machine reproduced every split, label, corruption assignment and count, with the default tree model matching the ledgered T1, T2 and T3 to every printed digit.

§2026.3(vii) Limitations

The corpus is a forward-model output, so these scores bound the pipeline's behaviour only within the declared physics; a real film can differ from the model in ways the corpus cannot show. The corruption catalogue is finite. The sp³ result depends on the presence of process metadata that a measured campaign may not carry. None of the figures above is a measurement, and none should be quoted as one.

How to cite

R. J. York, “Learned inversion of modelled carbon Raman spectra, with a refusal gate”, Finding 2026.3, ryanjamesyork.com, version 1.1, 2026-09-05. https://ryanjamesyork.com/findings/2026/3

@misc{york2026_3,
  author  = {York, Ryan James},
  title   = {Learned inversion of modelled carbon Raman spectra, with a refusal gate},
  year    = {2026},
  month   = {9},
  note    = {Finding 2026.3, version 1.1},
  url     = {https://ryanjamesyork.com/findings/2026/3},
  howpublished = {ryanjamesyork.com}
}
Compiled edition

This finding is compiled, with Finding 2026.2 and Finding 2026.4, in the preprint “Resolving the Raman crystallite size turnover in nanocrystalline graphite with a synthetic benchmark” (v2.0, 2026-09-05, PDF, 17 pp.). See Publications.

Provenance

Source: /findings/2026/3.md · SHA-256 17ed9da4c7a8361f2d1d315275f76ec6fa469047300432a347c44539c46c2de0
How to verify