TOPICS API LAB / API FOUNDATIONS
Topics API design: IDs, labels, and useful JSON contracts
Build a topics API contract that readers, editors, and developers can understand. Explore stable identifiers, evidence, uncertainty, document revisions, and vocabulary changes through practical examples for content classification.

A useful topics API gives a document a stable place in a larger collection. It lets a search interface filter by subject, a publisher build a focused reading list, and an analyst compare coverage across time. The difficult work begins before the first endpoint: deciding what a topic means, how an assignment is justified, and which parts of the contract can change.
This guide uses “topics API” to mean an interface for classifying content against a defined vocabulary. Google's browser Topics API documentation describes a separate interest-based advertising mechanism. The examples here concern document subjects, rather than browser interest signals. Start with the broader Topics API foundations if you are choosing the scope of your own system.
Define the decision before designing the payload
Write one sentence describing what the classification will enable. “Route incoming articles to editorial queues” leads to different choices from “support a public archive of long-term subjects.” The queue may accept broad, provisional labels because an editor reviews every result. The archive needs consistent concepts, clear revision history, and rules for replacing outdated assignments.
Next, identify the unit being classified. A headline, full article, paragraph, transcript, and collection are different inputs. A headline about a council vote may omit the transport policy discussed in the body. If your contract accepts either headline-only or full-text inputs, record that distinction explicitly. A consumer should not have to infer evidence coverage from the number of returned topics.
Finally, decide whether you need one label or several. A story about a university's solar installation could reasonably concern education, energy, and construction. A single primary topic can support navigation, while secondary topics support discovery. Define how the primary choice is made so that “primary” does not become a convenient but unexplained ranking.
Separate identifiers, labels, and definitions
Give every concept an identifier whose meaning survives routine wording changes. A label is the wording a reader sees; an identifier is the value another system stores. If an editor changes “Solar power” to “Solar energy,” a stable identifier lets existing assignments remain connected without rewriting every document. Avoid embedding the entire category path inside the identifier if reorganization is likely.
A definition should describe the inclusion boundary. “Reporting substantially about electricity or heat derived from sunlight” is more actionable than “Everything about solar.” Add an exclusion note for a common confusion, such as a fictional company named Solar. Include a few representative examples, but do not let examples replace the definition: future documents will always find a new edge case.
Keep aliases separately. “Photovoltaics” might help searchers discover the concept, while the preferred label stays short. An alias can also be narrower than a topic, so document whether it is an exact synonym or simply a retrieval aid. Treating every helpful search term as a new topic creates duplicates that later complicate reporting.
Keep concepts apart from assignments
A topic record describes a concept. An assignment describes the claim that a particular document belongs to that concept. Combining them into one mutable object makes revision difficult: changing a topic label should not silently change the evidence or review status attached to a document. Keep the vocabulary and the classification results connected by identifiers.
The following hypothetical response illustrates that separation. The identifiers, release names, and values are invented for this guide; they describe a possible application contract.
{
"document_id": "article-solar-campus",
"document_revision": "r3",
"taxonomy_version": "2024-03",
"status": "review_required",
"assignments": [
{
"topic_id": "energy.solar",
"role": "primary",
"evidence": "The campus will install rooftop solar panels."
}
]
}
The document revision matters because an edited article may no longer support the same assignment. The taxonomy version identifies the definitions used at decision time. A review status tells consumers what they may do next. The evidence gives a reviewer something concrete to inspect. None of these fields alone establishes correctness; together, they make the decision easier to examine.
If you include scores, explain their origin and interpretation. A ranking value, a calibrated probability, and a model's self-reported certainty are different things. Name the field accordingly, document its range, and specify whether comparisons across topics or model versions are meaningful. The topic classification evaluation guide develops the practical checks behind those choices.
Make uncertainty and empty results explicit
An empty topic list can mean several things: no supported subject, unreadable input, a missing vocabulary, or an interrupted classifier. Give these outcomes distinct states. A successful classification with no relevant labels should remain distinguishable from a request that never reached a usable decision. Otherwise, downstream reports may count operational failures as evidence that a subject disappeared.
Consider a brief consisting only of “More details soon.” A sensible response might be an insufficient-context state with an empty assignment array. For an unrelated but complete document, the response might instead be outside-scope. These states lead to different actions: retrieve more content for the first, and accept that the vocabulary does not cover the second.
Write handling rules for conflicting evidence too. A headline may emphasize a sports event while the body primarily discusses a sponsorship dispute. Your application can prefer full-text evidence, request editorial review, or expose separate classifications for separate fields. The important design choice is to make that behavior intentional and reproducible.
Design versions around meaning
Keep the interface version separate from the taxonomy version. Changing an optional display field is different from splitting a broad topic into two narrower concepts. The response shape may remain identical while the classification meaning changes substantially. A client that only checks the endpoint version would miss that change unless the vocabulary release is recorded.
For a renamed topic, retain the identifier when the definition stays equivalent. For a split, create explicit successor concepts and preserve the old record for historical interpretation. An old assignment to “Urban mobility” cannot automatically reveal whether the document belongs to cycling infrastructure or public transport. Mark a suggested migration as a suggestion when evidence must be reviewed again.
Publish small migration examples alongside each release. Show an unchanged document before and after a label edit, then show a document that needs reassessment after a split. This approach is especially useful for news topic workflows, where archive continuity and today's editorial vocabulary must coexist.
Specify behavior at the edges
Document what happens when the same document revision is submitted twice. Decide whether the operation returns a stored result or creates a new classification attempt. Either approach can be useful, but consumers need an attempt identifier when results may differ. Preserve enough context to distinguish an intentional reclassification from an accidental duplicate delivery.
For batch inputs, return an outcome for each submitted item. One malformed document should not make every other result ambiguous. Keep the caller's document identifier attached to errors as well as successes, and document whether response order follows input order. A client should be able to reconcile a batch without guessing which omitted item failed.
Also define text limits, accepted languages, and normalization rules. If your application removes boilerplate or truncates a long article, record what happened. A topic assignment based on the opening paragraphs deserves a different review from one based on the complete text. Store only the evidence needed for your workflow, and choose retention deliberately.
Review the contract with a consumer
Walk through three sample documents with the person building the next interface. Include one straightforward match, one ambiguous case, and one valid empty result. Ask them to explain what their screen would show and what action follows. Missing fields often become obvious when someone tries to render a useful review card.
Then inspect a historical result after a vocabulary change. Can the consumer still recover the original definition and document revision? Can they distinguish a machine suggestion from an editor's accepted assignment? If structured generation will produce your responses, use the LLM extraction contract guide to separate parsing checks from topic evidence.
Conclusion: make the decision inspectable
A strong contract connects stable concepts to specific documents and explains the limits of each assignment. Start with identifiers, definitions, evidence, and explicit outcomes. Add complexity only when a consumer can explain why it is needed. The result is easier to integrate, easier to review, and easier to revise when your content or vocabulary grows.


