# Score a toy binary task forecast

Illustrative planning brief; no automatic product import.

Synthetic planning case: Two toy forecasts each assign 0.8 to an accepted proof being ready by its declared checkpoint. Observed labels are one for ready and zero for not ready. A comparison assigns 0.5 to both.

## Decision

Under the stated mean squared probability-error rule, the 0.8 pair scores 0.34 and the 0.5 pair 0.25. This two-case comparison does not establish calibrated probabilities or a preferred real forecasting model.

## Owned work

- Define the binary event
  - Owner role: Evaluation owner
  - Acceptance evidence: Ready means accepted by the specific checkpoint; labels are genuinely known.
- Keep probabilities and outcomes separate
  - Owner role: Recorder
  - Acceptance evidence: The 0.8 inputs were supplied before the outcomes, not fitted afterward.
- Interpret the sample score narrowly
  - Owner role: Reviewer
  - Acceptance evidence: Two cases compare the stated rule without a calibration or future-confidence claim.

## Workflow

1. Define the binary event. Check: Ready means accepted by the specific checkpoint; labels are genuinely known.
2. Keep probabilities and outcomes separate. Check: The 0.8 inputs were supplied before the outcomes, not fitted afterward.
3. Interpret the sample score narrowly. Check: Two cases compare the stated rule without a calibration or future-confidence claim.

## Judgment

No real probability model, sampling design or calibration evidence is supplied. Binary scoring uses the page’s own explicit mathematical rule.


## Filled manual planning note

Under the stated mean squared probability-error rule, the 0.8 pair scores 0.34 and the 0.5 pair 0.25. This two-case comparison does not establish calibrated probabilities or a preferred real forecasting model. For each case use (p−y)². Mean at 0.8=(0.04+0.64)/2=0.34; at 0.5=(0.25+0.25)/2=0.25. Lower means less error under this explicit toy score only. No real probability model, sampling design or calibration evidence is supplied. Binary scoring uses the page’s own explicit mathematical rule.


## Workflow questions

### Does 0.8 mean this example proves an 80% success chance?

No. It is a supplied toy forecast value; this worksheet does not validate it.

### Can unknown outcomes be scored as zero?

No. Zero is an observed not-ready label under this definition; unknown remains unresolved.

## Product connection

Use the owned checks and downloaded brief to discuss this planning decision alongside your TeamBoostAI tasks. Confirm available fields, roles and account features separately. The example is manual; it does not calculate live analytics, create work or run an experiment in the product.

Confirm account availability before adopting this manual outline.

## Original worked case

Synthetic records, manual planning only. No account import or live analytics.

### Inspect the invented case records

Case | Supplied probability | Binary outcome | Squared error at 0.8 | Squared error at 0.5
--- | --- | --- | --- | ---
A | 0.8 | 1 | 0.04 | 0.25
B | 0.8 | 0 | 0.64 | 0.25

### Reasoning

For each case use (p−y)². Mean at 0.8=(0.04+0.64)/2=0.34; at 0.5=(0.25+0.25)/2=0.25. Lower means less error under this explicit toy score only.

### Bounded result

Under the stated mean squared probability-error rule, the 0.8 pair scores 0.34 and the 0.5 pair 0.25. This two-case comparison does not establish calibrated probabilities or a preferred real forecasting model.

### Distinct decision

This calculates a probability-error score, rather than comparing duration errors or promising task completion confidence.

### Limits

No real probability model, sampling design or calibration evidence is supplied. Binary scoring uses the page’s own explicit mathematical rule.

### Definitions and method context

- Forecasting: Principles and Practice — accuracy — https://otexts.com/fpp3/accuracy.html — Genuine held-out forecast and point-error evaluation context. Binary scoring examples use their own explicit toy definitions; no real task model is validated.
