# Freeze a forecast evaluation cutoff before fitting

Illustrative planning brief; no automatic product import.

Synthetic planning case: A manual evaluation declares days 1–5 for fitting and days 6–8 held out. A proposed parameter choice uses day-seven outcomes before claiming performance on days 6–8.

## Decision

The proposed rule has seen part of its held-out outcomes, so the declared evaluation is contaminated. Preserve the cutoff and choose a genuinely unseen evaluation or an explicitly different scope.

## Owned work

- Freeze fitting and test identities
  - Owner role: Evaluation planner
  - Acceptance evidence: The days 1–5 versus 6–8 split is recorded before rule selection.
- Inspect every information use
  - Owner role: Reviewer
  - Acceptance evidence: Day-seven outcome influenced parameter choice rather than being unseen evaluation evidence.
- Restore a valid declared evaluation
  - Owner role: Method owner
  - Acceptance evidence: A new untouched cohort or transparently revised claim is chosen; contamination is not erased from history.

## Workflow

1. Freeze fitting and test identities. Check: The days 1–5 versus 6–8 split is recorded before rule selection.
2. Inspect every information use. Check: Day-seven outcome influenced parameter choice rather than being unseen evaluation evidence.
3. Restore a valid declared evaluation. Check: A new untouched cohort or transparently revised claim is chosen; contamination is not erased from history.

## Judgment

The cutoff alone does not validate sample size, features, timing or forecasting performance.


## Filled manual planning note

The proposed rule has seen part of its held-out outcomes, so the declared evaluation is contaminated. Preserve the cutoff and choose a genuinely unseen evaluation or an explicitly different scope. The three declared held-out days include day seven. Using that outcome to select the rule violates the stated information boundary even if the final reported score uses all three rows. The cutoff alone does not validate sample size, features, timing or forecasting performance.


## Workflow questions

### Can held-out data be examined eventually?

Yes for evaluation; using it for tuning changes what that same cohort can support.

### Does this require a specific train/test percentage?

No. This example checks information boundaries, not a universal split ratio.

## Product connection

Use the owned checks and downloaded brief to discuss this planning decision alongside your TeamBoostAI tasks. Confirm available fields, roles and account features separately. The example is manual; it does not calculate live analytics, create work or run an experiment in the product.

Confirm account availability before adopting this manual outline.

## Original worked case

Synthetic records, manual planning only. No account import or live analytics.

### Inspect the invented case records

Evidence window | Declared use | Proposed actual use
--- | --- | ---
Days 1–5 | Fit rule | Fit rule
Day 6 | Held-out outcome | Evaluate
Day 7 | Held-out outcome | Used to choose parameter
Day 8 | Held-out outcome | Evaluate

### Reasoning

The three declared held-out days include day seven. Using that outcome to select the rule violates the stated information boundary even if the final reported score uses all three rows.

### Bounded result

The proposed rule has seen part of its held-out outcomes, so the declared evaluation is contaminated. Preserve the cutoff and choose a genuinely unseen evaluation or an explicitly different scope.

### Distinct decision

The artifact checks outcome leakage into fitting, separate from retaining a forecast vintage or rolling origins.

### Limits

The cutoff alone does not validate sample size, features, timing or forecasting performance.

### Definitions and method context

- Forecasting: Principles and Practice — accuracy — https://otexts.com/fpp3/accuracy.html — Genuine held-out forecast and point-error evaluation context. Binary scoring examples use their own explicit toy definitions; no real task model is validated.
