LLM Topics API
Turn language into structured, validated topic records.
TOPIC GUIDE / 02
AI topic classification connects unstructured text to a defined vocabulary. Explore label design, representative evaluation sets, and review policies that keep the output useful when the text gets complicated.

Decide what the system should classify: a sentence, a whole article, a transcript segment, or a collection of documents. The choice changes what context is available. A short headline may imply a subject that the full article only mentions in passing.
Write a definition and several contrasting examples for each category. Include a boundary example where the correct decision is debatable. Those cases often expose an unclear policy before any model is involved. Editors and developers should be able to explain the same label in similar terms.
A simple keyword rule can be a useful starting point for a narrow, well-defined category. A trained classifier or language model may recognize wording that the rule misses, but it also introduces new failure modes. Compare approaches on the same documents and the same label definitions.
Record the model or rule version separately from the taxonomy version. A change to the decision engine is different from a change to the meaning of a topic. Keeping both visible helps you investigate whether a shift in assignments reflects content, model behavior, or editorial policy.
Precision asks how many assigned labels are correct; recall asks how many relevant labels were found. Those questions serve different workflows. An automatic public label may need stricter review than a suggestion shown privately to an editor.
Inspect results by topic, language, document length, and source type. A single average can hide a category that never receives a correct assignment. Keep a held-out set that is separate from examples used to adjust prompts or rules, and record the decisions made when reviewers disagree.
Let uncertain documents move into an explicit review state. Preserve the text excerpt that supports an assignment when your data permissions permit it, and make the reason easy to inspect. An explanation should point to the document rather than simply repeat the category name.
Monitor a sample of accepted assignments after launch. Changes in writing style or subject matter can make earlier evaluation results less representative. Use those observations to update the review set deliberately, with a record of what changed and why.
This example highlights fields worth discussing when you design your own contract. Define their meanings, allowed values, and review rules before an application relies on them.
| FIELD | PURPOSE |
|---|---|
model_version | Decision engine identifier |
topic_ids | Allowed assigned concepts |
evidence | Supporting document excerpt |
review_status | Accepted or awaiting review |
{
"document_id": "example-brief",
"model_version": "example-classifier-v1",
"topic_ids": [
"ai.evaluation"
],
"evidence": "The team compared label errors.",
"review_status": "needs_review"
}Illustrative schema and example values; adapt them to your data and review process.
No. Treat any score according to how it was defined and evaluated. Compare scores with observed outcomes before using them to automate a decision.
Yes, if the task calls for multiple subjects. Define how secondary labels are selected and avoid assigning every concept that receives a brief mention.
Use representative and difficult examples, including overlapping labels, unfamiliar wording, multilingual text, and documents outside the vocabulary.
CONNECTED TOPICS
Turn language into structured, validated topic records.
Write clear instructions for consistent topic assignments.
Organize model tasks, versions, and capability evidence.