TOPICS API LAB / AI & LLMS
LLM topic extraction: a JSON contract you can actually validate
Valid JSON is the first check in a longer workflow. Design topic extraction around an allowed vocabulary, verifiable evidence, explicit outcomes, document revisions, and validation that connects a generated label to its source.

A language model can suggest useful topic labels, but an application needs a more precise agreement about what those suggestions mean. The response must identify the document, choose from the permitted vocabulary, and expose enough evidence for review. A neatly formatted object is useful only when the next component can decide whether to accept it.
The LLM Topics API guide introduces the broader extraction workflow. Here, the focus is the boundary between generated output and application data: a hypothetical JSON contract that can be parsed, checked, reviewed, and revised. The examples describe a design pattern for your own application, with no assumption that a particular model or hosted endpoint provides it automatically.
Choose the task and vocabulary first
Decide whether the model should discover candidate subjects or assign existing concepts. Discovery can suggest new wording for later editorial review. Classification should select from a known vocabulary. Mixing the two tasks in one unrestricted label array makes it difficult to distinguish a legitimate topic from a plausible phrase the model invented.
For classification, supply stable topic identifiers alongside concise definitions and exclusions. A display label by itself can be ambiguous. “Models” might refer to machine learning systems, statistical abstractions, or people in fashion photography. The definition should establish the intended meaning, while boundary examples clarify the cases most likely to cause confusion.
Also define how much evidence makes a topic relevant. Require substantive coverage if the result will populate subject pages. Permit weaker candidate matches only when a reviewer will inspect them before use. Make the primary-topic rule explicit, and set a reasonable maximum number of assignments. The Topics API schema guide explains these contract choices in more detail.
Describe the smallest useful result
Begin with document identity, a revision reference, an outcome, and an assignment list. Each assignment needs a permitted topic identifier and supporting evidence. Add fields only when a consumer can explain how it will use them. A long response packed with free-form analysis can be harder to validate than a small result with a few well-defined obligations.
The following invented schema illustrates a single assignment. It is intentionally limited to two sample topics so that the permitted values are easy to inspect.
{
"type": "object",
"properties": {
"topic_id": {
"type": "string",
"enum": ["transport.public", "energy.solar"]
},
"evidence": {"type": "string"},
"review_required": {"type": "boolean"}
},
"required": ["topic_id", "evidence", "review_required"],
"additionalProperties": false
}
The official JSON Schema object reference explains that naming properties does not make them required. The required list specifies mandatory fields, while setting additional properties to false rejects undeclared fields in this simple object. It also distinguishes a missing property from a property whose value is null. Decide which representation each field accepts rather than treating them as interchangeable.
This schema checks shape and allowed identifiers. It does not establish that the evidence appears in the document or that the selected topic fits that evidence. Those are separate obligations for the surrounding application and review process.
Validate in several understandable stages
First, parse the output as JSON. If parsing fails, keep the failure visible instead of extracting whichever fragment happens to resemble an object. A permissive repair step can accidentally turn a partial or contradictory response into something that looks authoritative. If you choose to repair formatting, preserve the original output and label the repaired attempt.
Second, validate the parsed result against the chosen schema. Reject unknown identifiers, missing required fields, and values of the wrong type. Check nested objects as carefully as the outer envelope. For batches, attach every outcome to its input identifier so that a rejected result cannot shift the association of the remaining documents.
Third, apply application rules. Confirm that the taxonomy version exists, the document revision matches, the assignment limit is respected, and duplicate topic identifiers are handled consistently. A schema can express some of these constraints, but the application still needs access to the relevant vocabulary and document records.
Keep semantic review separate
Finally, inspect whether the selected topics are supported by the actual text. A structurally valid assignment of solar energy to an article about a company named Solar can still be wrong. Preserve this distinction in logs and evaluation reports: parsing failures, contract failures, and unsupported classifications call for different improvements.
Make the evidence independently checkable
Ask for a short source passage that supports each assignment. In the application, confirm that the passage occurs in the exact document revision used for classification. If your workflow permits normalized quotation matching, define the normalization rules explicitly. Otherwise, an apparently harmless punctuation change may hide a more substantial mismatch.
Offsets can help a review interface highlight the relevant passage, but define what an offset counts. Bytes, Unicode code points, and displayed characters are not interchangeable in every text. Choose a representation, test it with the languages you support, and keep the source text stable between extraction and rendering. An exact quoted span provides a useful additional check.
Require the evidence to support the chosen concept, not merely share a keyword. A sentence stating that a company does not operate solar installations should not justify a solar-energy assignment simply because the phrase appears. Include negation, quotations, historical background, and speculative statements in your review examples so that the evidence policy has clear boundaries.
Give uncertainty a valid output
A model needs an acceptable way to return no supported assignment. Distinguish a complete document outside the vocabulary from text that is too brief or incomplete to classify. If the application cannot represent these states, it implicitly pressures every input into a topic, even when the available evidence does not justify one.
Use explicit outcome names and document what follows from each. An insufficient-context result may trigger retrieval of the full article. An outside-scope result may be retained without an assignment. A review-required result may enter an editorial queue. A processing error should remain a processing error, with enough context to retry the correct input.
Keep source documents separate from application instructions. An article may contain quoted commands, code, or language that resembles a prompt. Treat that material as content to classify. Test whether your implementation preserves the boundary, and validate output independently; wording the boundary clearly is one part of the design, not proof that every result will obey it.
Handle long documents and repeated attempts deliberately
If you divide long documents into sections, decide how section-level assignments become a document-level result. Repetition in several chunks should not automatically make a subject primary. A long appendix may repeat a term more often than the main argument. Preserve section context and use a documented aggregation rule that reflects the article's purpose.
Record whether the input was shortened, cleaned, or translated before classification. A reviewer looking at the full article needs to know which text the model actually received. Retain a processing revision alongside the document revision when a change to extraction rules could affect the result.
Define a bounded retry policy. A formatting failure may justify another generation attempt, while missing source text requires a different remedy. Avoid repeatedly asking for a preferred label until it appears. Give each attempt an identifier and keep the selected result traceable to the input, prompt, vocabulary, and validation outcomes that produced it.
Evaluate the complete contract
Measure more than the share of responses that parse. Review unknown-label rejection, evidence matching, supported-topic accuracy, and the usefulness of uncertain outcomes. Test the complete path from document preparation to the final review screen. The classification evaluation guide shows how to connect these checks to representative documents and practical release decisions.
When changing the prompt or vocabulary, compare results on the same inputs. Inspect changed assignments and their evidence, especially where a candidate becomes more specific. Use the Prompts Topics API guide to organize instructions and examples around the exact boundaries your evaluation identifies.
Conclusion: connect every label to a check
A useful extraction contract makes each obligation visible: valid structure, permitted concepts, current document identity, and supporting evidence. Let uncertainty remain explicit, and keep structural checks distinct from semantic review. That gives developers a dependable integration boundary and gives editors a clear reason to accept, revise, or reject each proposed topic.


