TOPICS API LAB / AI & LLMS
Evaluate AI topic classification before you trust the labels
A useful evaluation starts with the decisions a label will drive. Build a representative review set, calculate topic-level metrics, inspect disagreements, and compare revisions without hiding costly mistakes in a single average.

An AI topic classifier becomes useful when its labels support a specific decision. That decision might place an article in a reading list, send a document to an editor, or help a researcher find coverage of an emerging subject. Evaluation should ask whether the classifier supports that decision reliably enough, with errors that the team understands and can manage.
A single score cannot tell you whether a system confuses adjacent subjects, misses short documents, or confidently labels articles that fall outside its vocabulary. Begin with the AI Topics API overview, then use this guide to design a practical evaluation around your own content. Every numerical example below is hypothetical, and no result represents a TopicsAPI service benchmark.
Write the acceptance rule first
Before collecting examples, describe what counts as a correct assignment. A policy article mentioning a university is not necessarily about education. A story discussing an electric bus fleet might legitimately need both transport and energy labels. Reviewers need a shared rule for substantive coverage, passing mentions, and documents with several important subjects.
Separate the primary-topic decision from the presence of secondary topics. A classifier might identify all relevant subjects while choosing an unhelpful primary label for navigation. Score those behaviors separately if the interface uses them differently. Otherwise, improvements to secondary coverage could hide a decline in the main category that readers actually see.
Document acceptable uncertainty too. A short, incomplete brief may deserve a request for more context. An article outside the vocabulary may deserve no label. Treat those as intentional outcomes when your workflow allows them, while keeping input failures separate. The Topics API contract guide explains how to represent these distinctions without overloading an empty array.
Build a review set that resembles the work
Collect examples from the inputs the application will really receive. Include long features, short updates, mixed-subject articles, and text with normal formatting noise. If the production input contains a title and body, preserve both in the evaluation set. Testing clean summaries while processing raw article text answers a different question from the one your users face.
Record useful attributes alongside each document: language, publication format, length range, broad subject, and whether it has been revised. These fields allow later inspection without changing the reference labels. Keep the original text or a stable revision reference so that a future reviewer can reconstruct the actual input.
Choose samples from separate stories, rather than allowing many near-identical updates to dominate the set. Ten versions of one announcement should not count as ten independent demonstrations of broad coverage. Place closely related versions together when dividing development and evaluation material, and explain how you handled repeated or syndicated text.
Include deliberate boundary cases
Add documents that distinguish neighboring labels. For an energy taxonomy, that might include a solar installation report, a utility earnings story, and a profile of a company whose name contains “Solar.” These examples help expose the reasoning implied by the vocabulary. Keep their role visible: a challenge set tests known boundaries, while a representative sample estimates ordinary workflow behavior.
Create reference labels through review
Have reviewers apply the written definitions before looking at model output. Otherwise, a plausible machine suggestion can become an anchor for the reference answer. For important boundary cases, ask two people to label independently, then resolve disagreements by recording the evidence and the rule that decided the case.
A disagreement does not automatically mean one reviewer was careless. It may reveal an overlapping definition, an incomplete article, or an unresolved editorial policy. Preserve such cases in a disagreement log. If a definition changes, review the affected examples and create a new version of the reference set rather than silently editing yesterday's answer.
Record unsupported assignments as well as missing ones. A reviewer should be able to say that a document supports energy but does not support finance, even though a company name appears. This negative evidence makes recurring confusion easier to diagnose and provides useful material for improving topic-labeling prompts.
Measure precision and recall for each topic
Google's classification metrics guide defines precision as correct positive predictions divided by all positive predictions, and recall as correct positive predictions divided by all actual positives. Both describe performance at a chosen decision threshold. Overall accuracy can conceal poor behavior when a topic is rare, because correctly rejecting many irrelevant documents can dominate the total.
Suppose a hypothetical evaluation contains twenty articles that truly concern solar energy. A classifier assigns the solar label to eighteen articles: fifteen correct matches and three unrelated articles. It misses five relevant articles. Precision is fifteen divided by eighteen, approximately eighty-three percent. Recall is fifteen divided by twenty, seventy-five percent. These figures describe different consequences: unnecessary inclusions and missed coverage.
Show the counts beside the percentages. A result based on a handful of examples deserves different confidence from one supported by a large and varied sample. If a topic has no positive predictions or no reference positives, report the missing denominator explicitly and explain your reporting convention. Do not let a software default create an impressive but meaningless number.
Inspect the errors behind the metrics
Group mistakes by cause. Useful categories include passing-mention confusion, neighboring-topic confusion, unsupported specificity, missed secondary topics, and insufficient context. These categories connect a score to an action. A broad labeling definition needs different work from a text-extraction problem that removed the paragraph containing the main subject.
Look at the same categories across document slices. A classifier might work well on long English articles yet struggle with brief bilingual updates. You do not need a separate dashboard for every attribute, but you should examine the dimensions that matter to your users. Combine very small slices carefully and retain the underlying counts.
Review successful cases too. A correct label attached to irrelevant evidence may be a lucky result that will fail on the next document. Ask whether the cited passage actually supports the topic definition. Evidence review is particularly helpful in LLM topic extraction, where fluent explanations can make weak decisions look persuasive.
Tune decisions without moving the test
If your classifier exposes a useful score, experiment with decision thresholds on a development set. Decide how you will trade additional review work against missed documents. A public subject page might demand stronger evidence than an internal discovery queue, so the two destinations may need different acceptance rules even when they use the same underlying classifier.
Keep a separate evaluation set for the final comparison. Repeatedly inspecting its failures and rewriting prompts around them turns it into development material. When that happens, label it honestly and reserve fresh examples for the next release decision. Preserve the old set for regression checks without pretending it remains untouched.
Do not interpret a generated confidence number as a measured probability by default. Compare score ranges against observed correctness on reviewed examples, and document what the score actually represents. When evidence is thin, a clear review state may be more useful than a precise-looking decimal that nobody can defend.
Make release comparisons actionable
Compare candidate versions on identical document revisions and reference definitions. Record the model identifier, prompt revision, vocabulary release, preprocessing rules, and decision thresholds. Then review which individual assignments changed. An improved aggregate score can still introduce a serious new confusion in the subject that matters most to one editorial team.
Set acceptance criteria before choosing a winner. A hypothetical team might require fewer unsupported primary labels, no regression on its essential subject categories, and a review queue that fits the editors' capacity. Those are local product decisions, not universal thresholds. State their rationale so that the next team member can judge whether they still fit the workflow.
Conclusion: make evaluation a decision record
A useful evaluation connects definitions, examples, errors, and consequences. Keep the reference set versioned, show counts with metrics, and inspect the evidence behind both successes and failures. The goal is a clear account of where topic classification helps, where people still need to review it, and what the next change should improve.


