# Pregame Simulator Claim Boundary

Allowed wording:

- fantasy decision support
- probabilistic simulation
- pregame scenario estimate
- ITHKOR-SIM inspired adaptive simulation

Blocked wording:

- guaranteed prediction
- betting odds
- sure win
- AI knows the result
- proves ITHKOR
- physical claim

## Product Boundary

The simulator helps compare fantasy lineup scenarios before a game. It estimates
possible fantasy-point ranges from available mock/CSV/player context. It does
not predict the final match result.

## Research Boundary

SIM1, SIM2C and SIM3C are used as engineering heuristics:

- spend more compute on uncertain/high-impact players;
- keep pregame checkpoint records;
- store observable capsules for audit.

This does not claim that ITHKOR is proven, that the universe is a simulation,
or that the model has physical correctness beyond its fantasy scoring context.

## Data Boundary

Elite Prospects data must come from an approved API/license path. The module
must not scrape Elite Prospects pages without explicit permission.

The public prototype can use mock/CSV data for harness checks and
`wfm-public` data for the live public player pool while a licensed feed is not
available.

The `wfm-public` source reads public fantasyhockey.eu assets and anonymous JSON
endpoints. It may be described as a public WFM baseline for simulation plumbing.
It must not be described as authenticated WFM production history, licensed Elite
Prospects ingest, confirmed next-round lineup news, or model accuracy evidence.
Snapshot source URLs, asset versions and SHA-256 hashes may be used as
prediction provenance only; they prove which public inputs were read, not that a
prediction was correct.

The `nhl-public` source reads public `api-web.nhle.com` schedule, score,
standings and club-stats endpoints into a local snapshot. It may be described
as a real NHL model-lab source for roster/stat simulations and profile
experiments. It must not be described as direct WFM Slovak Extraliga production
evidence or as authenticated WFM postgame history.

NHL boxscore actuals exported from `api-web.nhle.com/v1/gamecenter/{gameId}/boxscore`
may be used to verify NHL model-lab recommendations and exercise the postgame
actuals pipeline. They must not be mixed into WFM production promotion gates or
labelled as `authenticated_wfm`.

NHL model-lab backtest cases generated from public NHL recommendations and
boxscore actuals must remain labelled with
`actuals_source_type=nhl_public_boxscore`. Their metrics may be used to debug
projection, interval and recommendation mechanics, but they are nonproduction
evidence even when imported into a temporary or explicit historical-case ledger.
Batch cases generated from the current NHL season snapshot are retrospective
diagnostics because the public season aggregates may include information after
the evaluated game date. They must not be described as no-lookahead historical
production backtests.

NHL model-lab sweep reports may be described as bounded retrospective
diagnostics over candidate profile parameters. Their `diagnostic_ready` gate
means only that the model-lab sample is large enough to inspect. It must not be
described as interval quality passing, automatic parameter promotion,
authenticated WFM evidence, or production readiness.

Public role estimates inferred from season production must remain labelled as
estimates.

## Backtest Boundary

The current historical backtest file contains synthetic fixture cases. It is
useful for proving that the prediction-verification pipeline, metrics and
profile comparison work end to end.

It must not be described as production model accuracy. Production accuracy can
only be claimed after real pregame snapshots and postgame WFM fantasy results
are imported, frozen, and evaluated on held-out cases.

In-sample tuning results must be described as diagnostics only. Model promotion
requires hold-out validation, enough verified cases, baseline improvement, and
credible interval coverage.

Backtest promotion also requires a passing evidence gate. Evaluated cases must
come from imported authenticated WFM postgame actuals with
`actuals_source_type=authenticated_wfm` plus provider, export id and export
timestamp metadata. Static fixture cases, mock results, public-only data and
CSV/manual imports without that provenance are nonproduction evidence even when
their error metrics look good.

Production promotion reviews must run with `evidence_scope=authenticated_wfm`.
The default all-case scope is a diagnostic harness and may include synthetic
fixtures; it must not be cited as production evidence.

Walk-forward results may be described as rolling out-of-time validation. They
must not be described as production accuracy while the input data is synthetic,
too small, or missing authenticated WFM postgame verification.

Interval calibration may be described as observed coverage quality on evaluated
cases. It must not be described as a guarantee that a future player or lineup
result will fall inside the interval.

Interval diagnostics may be described as a calibration review tool: residual
quantiles, error-to-width ratio and risk-band coverage on evaluated cases. They
must not be described as reliable future uncertainty estimates until they pass
the interval quality gate on enough authenticated WFM cases and the normal
monitoring/backtest gates also pass.
Calibration advice derived from those diagnostics may be described as a manual
review suggestion only. It must not be described as an automatic parameter
change, deployed model, or proof that future intervals are reliable.

## Lineup Recommendation Boundary

Lineup recommendations may be described as:

- ranked fantasy lineup scenarios
- strategy suggestions for expected, balanced or upside builds
- Monte Carlo-checked decision support

They must not be described as:

- guaranteed best lineup
- betting edge
- production accuracy proof
- evidence that the model knows the next result

The optimizer ranks candidates from the available projected player pool. Until
authenticated WFM lineup locks, current roles and verified postgame fantasy
outcomes are imported, recommendations are workflow and model-lab evidence only.

Locked recommendation snapshots may be described as pregame audit records.
Postgame recommendation verification may be described as measurement evidence
for those locked recommendations.

Batch round snapshots may be described as operational locking of multiple
pregame recommendation records, not as stronger accuracy evidence.

Scheduled pre-round capture may be described as an audit trail of what the
system recommended before a game. It must not be described as model accuracy
until final WFM points are imported and the monitoring/backtest gates pass.

Pending verification queues may be described as operational worklists for
collecting final fantasy points. They must not be described as model evidence
until the listed snapshots are actually verified.

Batch lineup verification may be described as grouped postgame measurement of
locked snapshots. It must not be described as production accuracy unless the
resulting verified cases later satisfy the normal readiness and held-out
backtest gates.

Recommendation verification templates may be described as operational CSV
artifacts for collecting final WFM fantasy points. A template, even when filled,
is not production accuracy evidence until it is verified against the locked
snapshot and evaluated through the normal historical/backtest gates.

Player-level actuals CSV filling may be described as a deterministic conversion
from supplied postgame player fantasy points to lineup-level
`actual_lineup_points`. It must not be described as authenticated evidence
unless the player-level source itself is authenticated and preserved.

They must not be described as model promotion, production accuracy, or proof of
predictive edge unless the exported cases later pass the normal held-out
backtest promotion gate with enough verified WFM data.

The recommendation report/readiness gate may be described as monitoring
evidence. It must not be described as promotion or production readiness when it
is based on mock, CSV, public snapshot data, too few recommendations, or
implausible interval coverage. A failed `interval_*` readiness reason means the
intervals are still diagnostic and should not be used as production confidence
bands.

The cycle audit may be described as an operations checklist for pregame capture,
pending postgame actuals, scorecard readiness and tuning evidence. It must not
be described as production prediction quality, an automatic tuning decision, or
proof that a competition profile is ready. Its gates summarize whether required
evidence exists; they do not create the evidence.

The runbook may be described as read-only operator guidance generated from the
cycle audit. It must not be described as an executed capture, verified postgame
import, scheduler, or automatic historical-case append. Any payload placeholder
such as `<contents of filled.csv>` still requires a real authenticated WFM
postgame source before it becomes evidence.

Prediction packages may be described as pregame handoff artifacts that lock
lineup recommendation snapshots and bundle pending verification CSV, provider
hashes, a manifest template and an operator checklist. They must not be
described as verified outcomes, production accuracy, postgame evidence or a
model promotion. A package only becomes useful evidence after authenticated WFM
postgame actuals fill the manifest, intake passes the hash/provenance gates,
verification rows are appended and the normal monitoring/backtest gates pass.

The postgame intake may be described as a manifest-checked preparation step for
authenticated WFM player actuals. It must not be described as verified accuracy
or historical evidence until its generated payload is submitted to the
verification endpoint, appended successfully, and then evaluated through the
normal monitoring/backtest gates. A manifest with a public, mock, missing, or
hash-mismatched source remains blocked evidence.

When appended, the manifest provenance may be described as the postgame actuals
source for verification rows and imported historical cases. This is distinct
from the pregame prediction snapshot source; a public pregame input does not
make authenticated postgame actuals public evidence, and authenticated actuals
do not by themselves prove prediction accuracy.

Competition scorecards may be described as per-league or per-profile monitoring
views over verified recommendations. They must not be described as automatic
profile tuning, production readiness, or a cross-league proof of accuracy unless
the underlying rows come from authenticated WFM postgame evidence and the
normal monitoring/backtest gates pass. Their source gate is based on postgame
verification provenance, not merely the pregame snapshot provider. A mixed
scorecard that contains mock, CSV, or unauthenticated verification rows remains
nonproduction evidence even when another profile in the same report has
authenticated rows.

Competition readiness reports may be described as per-competition model-ops
summaries that gather scorecard, interval, imported-case, backtest and tuning
proposal gates into a status and next action. They must not be described as an
executed verification, an automatic profile promotion, deployed parameters, or
production accuracy. A `ready_for_manual_profile_review` status still requires
human review and intentional parameter deployment; any failed gate keeps the
competition in evidence-building mode.

Model iteration reports may be described as ranked model-ops queues across
competitions. They may compare evidence gaps, calibration advice, candidate
patches and review URLs. They must not be described as automatic tuning,
deployed parameters, production accuracy, or proof that a competition is ready;
they only organize the next manual action after the normal evidence gates.

Profile shadow trials may be described as pregame current-vs-candidate
comparisons over the same target games. The candidate patch is request-local
and may be used to compare expected points, interval width, risk, lineup
overlap and captain changes before the round is played. It must not be
described as a deployed profile, automatic parameter promotion, verified
accuracy evidence, or production confidence until locked predictions are
verified with authenticated WFM postgame actuals and the normal gates pass.

Profile shadow packages may be described as locked paired pregame measurement
records for current and candidate profiles. They may include a pending CSV,
manifest, summary and checklist for later postgame verification. They must not
be described as candidate deployment, automatic profile promotion, production
accuracy, or historical training evidence. Candidate shadow rows should be
verified with `append_historical_cases=false` unless and until a human review
intentionally registers/promotes that candidate profile.

Profile shadow reports may be described as postgame current-vs-candidate
measurement over locked shadow pairs. They may compare MAE, RMSE, bias,
coverage and same-export authenticated WFM provenance. They must not be
described as automatic promotion, deployed parameters, or production accuracy;
they only support manual review after the normal evidence gates.

## Tuning Evidence Boundary

Stored backtest results may be described as an audit trail for model-tuning
decisions. They must not be described as a promoted model, production accuracy
claim, or proof of predictive edge unless the promotion gate passes on verified
held-out WFM cases.

Expanded optimizer runs may be described as a bounded comparison of candidate
parameter grids by competition. They must not be described as an automatic
profile promotion, especially when the input cases are synthetic or the
promotion gate fails.

Profile tuning proposals may be described as review objects that compare a
candidate profile with the currently registered profile. They must not be
described as deployed parameters, production readiness, or production accuracy
unless the proposal gate and normal backtest promotion gate pass on enough
imported verified WFM cases and a profile change is intentionally reviewed.

Synthetic stored runs are useful for proving the ledger and comparison workflow
only.

## Historical Import Boundary

Imported WFM historical cases may be described as verified measurement inputs
only after they include a known competition profile, locked pregame timestamp,
six-player lineup, captain, game ID, and final fantasy lineup points.

An imported case does not prove production accuracy by itself. It becomes part
of the evidence base only through backtesting, hold-out validation, baseline
comparison, and interval calibration.

Duplicate import prevention is a data-quality control, not an accuracy claim.
The synthetic smoke dataset and imported WFM cases must stay distinguishable in
status/debug output.

## Snapshot Boundary

Prediction snapshots are locked audit records. They may be described as:

- pregame forecast snapshot
- postgame verification metric
- model monitoring evidence

They must not be described as:

- guaranteed forecast
- betting edge
- proof that the next game result is known

Snapshot verification proves only how the locked pregame estimate compared with
the imported fantasy result for that case.

Appending verification exports into the historical case ledger is a data
lineage step. It may be described as closing the measurement loop, but not as a
model promotion. Promotion still requires enough verified cases, hold-out
validation, baseline improvement and credible interval calibration.
