TOPICS API LAB / APPLIED TOPICS
Design Prediction Topic Contracts Without Inventing Forecasts
Prediction coverage needs careful boundaries between the subject, the statement, and its time horizon. Build a document contract that preserves uncertainty and provenance without presenting extracted labels as forecasts or probabilities.

A document about the future can contain a forecast, a target, a scenario, a conditional statement, and a question in the same paragraph. A topic classifier that labels all of them prediction loses the differences that matter most to a reader. The first design task is to describe what the document says without adding a prediction of your own.
This guide develops a hypothetical publishing contract for future oriented coverage. It describes document metadata and editorial review, not a forecasting service. The Prediction Topics API overview introduces the subject area; the contract below shows how a team could make its records precise enough for search, comparison, and later correction.
Start with the statement being represented
Separate the document topic from any statement extracted from it. A topic might be library digitization. A statement might describe a proposed completion date, a conditional scenario, or an author's expectation. The topic remains useful even when the document contains no quantified forecast at all. Keeping the two records distinct prevents an empty probability field from looking like a missing product feature.
Use statement types that editors can identify from the source. A target expresses an intended outcome. A scenario describes a set of assumptions and possible consequences. A forecast expresses an expectation attributed to someone. A question asks about an outcome without asserting it. A historical report describes what happened, even if it quotes an older forecast.
In a hypothetical article, a library says it aims to digitize an archive within two years, subject to funding. Represent that as a stated target with a funding condition. Do not convert the target into a probability of completion, and do not treat the classifier's confidence in the label as confidence in the outcome.
Give time several separate fields
At least three times may matter: when the source was published, when the statement was made, and the period the statement concerns. A fourth field can record when your system retrieved the source. These times can differ. A new article may quote an older presentation, and a retrieved page may have changed after its original publication.
The W3C Time Ontology in OWL describes instants, intervals, durations, and relationships among temporal entities. For an editorial contract, this is a helpful conceptual basis for distinguishing a statement date from an outcome interval. The example fields here are a proposed application design, not an implementation of the complete ontology.
Store the original time phrase alongside any normalized interpretation. If a speaker says next summer, keep that wording and identify the reference date used to interpret it. If the source does not provide enough context, preserve an unresolved horizon. A precise looking date created from an ambiguous phrase can be more misleading than a visibly incomplete record.
Preserve granularity
A year, a quarter, and a day carry different precision. Do not invent midnight timestamps for records that only identify a year. When intervals are useful, document whether their endpoints are inclusive and which calendar or timezone applies. Consumers should be able to distinguish a source's precision from convenience added by the storage system.
Keep uncertainty attached to its owner
There are several different uncertainties in this workflow. The source may be uncertain about an outcome. The extractor may be uncertain about the statement type. An editor may be uncertain about the identity of the speaker. Treat these as separate observations with separate owners. A single confidence number cannot explain them all.
If a source supplies a probability, preserve its wording, scope, and method when available. Record that the value is reported by the source. If no probability appears, use an explicit absent state rather than generating a number. Qualitative expressions such as possible or likely can remain text unless your publication has a documented, appropriate mapping supplied by that source.
For the hypothetical library article, the record might contain an unresolved horizon, a target statement type, and a funding condition. That is already useful metadata. Readers can find conditional plans without being shown an invented confidence score. The extractor can separately indicate that the target classification requires human review because the paragraph mixes intention and expectation.
Design the contract around inspectable evidence
A compact record can include a stable statement identifier, its parent document, subject identifiers, statement type, exact source span, attributed speaker, temporal fields, conditions, and review state. Add a schema version and extraction version so later changes can be traced. Store identifiers independently of display labels to support renaming without breaking links.
The topic schema guide explains the broader value of stable identifiers and explicit contract boundaries. Here, those boundaries matter because downstream applications may otherwise interpret a richly structured statement as an endorsed conclusion.
- Subject: the thing being discussed, such as the hypothetical archive digitization project.
- Claim context: who made the statement, where it appears, and the conditions included in the source.
- Horizon: the original phrase, any supported normalization, and unresolved ambiguities.
- Review: extraction status, editorial decisions, and the reason for a correction.
Keep explanatory notes readable. A reviewer should be able to see why a record is unresolved without reverse engineering an internal error code. Structured fields serve retrieval; concise notes preserve the reasoning behind difficult decisions.
Distinguish information that is not stated from information that is not applicable or unresolved. A target with no probability is different from a damaged source passage whose probability cannot be read. Give each missing field a reason when the distinction affects interpretation. Clients can then display a useful explanation, avoid unnecessary retries, and route only the records that need editorial attention. This also makes completeness metrics more honest: an intentionally absent value is not an extraction failure.
Handle revisions without erasing the past
A statement may be revised, withdrawn, clarified, or discussed after its horizon has passed. Preserve the original record and connect the newer evidence through an explicit relationship. An updated source date does not necessarily mean the statement changed. Compare the relevant passage before changing its meaning or lifecycle state.
Consider a fictional follow up article that narrows the library project to one collection. That could revise the scope of the earlier target. It could also describe a separate phase. An editor needs both source passages to decide. Automatically treating every later mention as a replacement would destroy useful context.
Use neutral lifecycle labels such as active coverage, revised statement, or awaiting outcome review. An elapsed horizon alone does not establish whether an outcome occurred. If a publication wants to evaluate outcomes, define a separate evidence process and keep its conclusions distinct from the original extraction.
Test ambiguity before scale
Create a review set with clear targets, explicit scenarios, quoted forecasts, rhetorical questions, and retrospective reporting. Include statements with several subjects or horizons. Ask reviewers to identify both the record type and the evidence span. Agreement on a broad topic can hide disagreement about the sentence that supposedly supports it.
Measure errors that affect interpretation: assigning a probability that is absent, dropping a condition, attaching a quote to the wrong speaker, or normalizing a horizon from the wrong reference date. These failures deserve more attention than minor differences in display wording. The structured extraction guide offers related ideas for validating records before an application uses them.
Keep an abstention path. If a paragraph combines several incompatible interpretations, a review queue is a valid result. A system that always emits a complete object may appear reliable while filling its most consequential fields with assumptions.
Make the reading interface match the contract
Display the statement type beside the source attribution. Put conditions close to the statement and show unresolved timing in ordinary language. Avoid a dashboard tile that collapses a target, a scenario, and a forecast into one apparent outcome metric. The visual hierarchy should help readers understand what was said and by whom.
Allow filtering by subject and horizon without implying that matching records are comparable predictions. Two documents can discuss the same project while referring to different phases, assumptions, or definitions of completion. The record view should make those differences discoverable before inviting comparison.
Conclusion: describe future oriented coverage faithfully
A reliable prediction topic contract preserves subjects, attribution, time, conditions, and uncertainty. It can make future oriented documents searchable without asserting that their outcomes will occur. Start with explicit statement types, retain the original evidence, and treat unresolved fields as meaningful information. That produces a useful editorial dataset whose limits remain visible to every consumer.


