# Compare forecast methods on matched cases

Illustrative planning brief; no automatic product import.

Synthetic planning case: Two methods predict the same four case durations in minutes. A’s absolute errors are 2,4,8,10. B has errors 1,3 for the first two cases, but its last two predictions are missing. Their reported MAEs are six and two.

## Decision

The six-versus-two comparison uses different case sets. On the two matched cases A’s MAE is three and B’s two; retain that narrower observed comparison while leaving the four-case B result unavailable.

## Owned work

- Freeze target, unit and case identities
  - Owner role: Evaluation owner
  - Acceptance evidence: Both methods target the same duration in minutes on cases 1–4.
- Build the matched evidence set
  - Owner role: Reviewer
  - Acceptance evidence: Only cases 1 and 2 have errors for both methods; missing predictions remain visible.
- Bound the comparative claim
  - Owner role: Planning lead
  - Acceptance evidence: B’s smaller observed MAE is stated for the two matched cases, not all four or future cases.

## Workflow

1. Freeze target, unit and case identities. Check: Both methods target the same duration in minutes on cases 1–4.
2. Build the matched evidence set. Check: Only cases 1 and 2 have errors for both methods; missing predictions remain visible.
3. Bound the comparative claim. Check: B’s smaller observed MAE is stated for the two matched cases, not all four or future cases.

## Judgment

All errors are invented and already use the same unit. The matched subset gives a descriptive comparison only, with no imputation or model validation.


## Filled manual planning note

The six-versus-two comparison uses different case sets. On the two matched cases A’s MAE is three and B’s two; retain that narrower observed comparison while leaving the four-case B result unavailable. A all-case MAE=(2+4+8+10)/4=6. B observed-case MAE=(1+3)/2=2. Matched-case A MAE=(2+4)/2=3, versus B’s 2 on exactly those same cases. Missing predictions prevent a four-case B MAE; no zero error is imputed. All errors are invented and already use the same unit. The matched subset gives a descriptive comparison only, with no imputation or model validation.


## Before / after

Before: A six-minute all-case MAE ranked against B two-minute observed-case MAE.

After: Matched-case MAEs three and two reported beside the missing two-case B coverage.


## Workflow questions

### Can I hide cases 3 and 4 after pairing?

No. Preserve the excluded identities and the missing-prediction reason alongside the narrower comparison.

### Does matching two cases remove selection bias?

No. Missing predictions may be informative; the small paired subset cannot validate general superiority.

## Product connection

Use the owned checks and downloaded brief to discuss this planning decision alongside your TeamBoostAI tasks. Confirm available fields, roles and account features separately. The example is manual; it does not calculate live analytics, create work or run an experiment in the product.

Confirm account availability before adopting this manual outline.

## Original worked case

Synthetic records, manual planning only. No account import or live analytics.

### Inspect the invented case records

Case | A absolute error in minutes | B absolute error in minutes | Matched comparison
--- | --- | --- | ---
1 | 2 | 1 | Eligible pair
2 | 4 | 3 | Eligible pair
3 | 8 | Missing prediction | No pair
4 | 10 | Missing prediction | No pair

### Reasoning

A all-case MAE=(2+4+8+10)/4=6. B observed-case MAE=(1+3)/2=2. Matched-case A MAE=(2+4)/2=3, versus B’s 2 on exactly those same cases. Missing predictions prevent a four-case B MAE; no zero error is imputed.

### Bounded result

The six-versus-two comparison uses different case sets. On the two matched cases A’s MAE is three and B’s two; retain that narrower observed comparison while leaving the four-case B result unavailable.

### Distinct decision

The artifact exercises matched-case comparison after a misleading same-metric aggregate ranking, rather than converting hours to minutes.

### Limits

All errors are invented and already use the same unit. The matched subset gives a descriptive comparison only, with no imputation or model validation.

### Definitions and method context

- Forecasting: Principles and Practice — accuracy — https://otexts.com/fpp3/accuracy.html — Genuine held-out forecast and point-error evaluation context. Binary scoring examples use their own explicit toy definitions; no real task model is validated.
