# Keep unresolved outcomes outside scored verdicts

Illustrative planning brief; no automatic product import.

Synthetic planning case: A four-case binary evaluation has one correct prediction, one incorrect prediction and two unresolved outcome labels. A draft reports one success out of four.

## Decision

Known-label accuracy is one of two, or 50%, with two of four labels unresolved. Do not silently count missing outcomes as failures or generalize the known subset to all cases.

## Owned work

- Define the evaluation-label rule
  - Owner role: Evaluation owner
  - Acceptance evidence: A scored outcome must be observed under the agreed binary event definition.
- Keep unresolved cases in the ledger
  - Owner role: Recorder
  - Acceptance evidence: C and D are not removed from visibility or relabelled failed.
- Publish denominator and coverage together
  - Owner role: Reviewer
  - Acceptance evidence: The one-of-two result and two missing labels accompany the interpretation.

## Workflow

1. Define the evaluation-label rule. Check: A scored outcome must be observed under the agreed binary event definition.
2. Keep unresolved cases in the ledger. Check: C and D are not removed from visibility or relabelled failed.
3. Publish denominator and coverage together. Check: The one-of-two result and two missing labels accompany the interpretation.

## Judgment

This exercise supplies no missing-data model or population accuracy estimate. Unknown is an evidence state, not a negative outcome.


## Filled manual planning note

Known-label accuracy is one of two, or 50%, with two of four labels unresolved. Do not silently count missing outcomes as failures or generalize the known subset to all cases. Known scored cases=1+1=2. Known-label accuracy=1/2=50%; label availability=2/4=50%. One-of-four would assert failures for both unknown labels. This exercise supplies no missing-data model or population accuracy estimate. Unknown is an evidence state, not a negative outcome.


## Workflow questions

### Can missingness be systematic?

Yes. Harder or slower cases may be unresolved more often; the scored subset may not describe all forecasts.

### Should late labels replace the frozen score silently?

No. Issue an explicit evaluation update with its label cutoff and changed cases.

## Product connection

Use the owned checks and downloaded brief to discuss this planning decision alongside your TeamBoostAI tasks. Confirm available fields, roles and account features separately. The example is manual; it does not calculate live analytics, create work or run an experiment in the product.

Confirm account availability before adopting this manual outline.

## Original worked case

Synthetic records, manual planning only. No account import or live analytics.

### Inspect the invented case records

Case | Outcome evidence | Score treatment
--- | --- | ---
A | Known: prediction correct | Correct
B | Known: prediction incorrect | Incorrect
C | Unknown | Unresolved, not scored
D | Unknown | Unresolved, not scored

### Reasoning

Known scored cases=1+1=2. Known-label accuracy=1/2=50%; label availability=2/4=50%. One-of-four would assert failures for both unknown labels.

### Bounded result

Known-label accuracy is one of two, or 50%, with two of four labels unresolved. Do not silently count missing outcomes as failures or generalize the known subset to all cases.

### Distinct decision

The artifact distinguishes forecast-label availability from accuracy among observed labels.

### Limits

This exercise supplies no missing-data model or population accuracy estimate. Unknown is an evidence state, not a negative outcome.

### Definitions and method context

- Forecasting: Principles and Practice — accuracy — https://otexts.com/fpp3/accuracy.html — Genuine held-out forecast and point-error evaluation context. Binary scoring examples use their own explicit toy definitions; no real task model is validated.
