Search Methodology
The dashboard stores only primary sources in the source jurisdiction's own language: official documents, speeches, reports, and publications authored by tracked institutions and actors. When an English-language secondary source references a primary document, the pipeline preserves the secondary source in the entry's Sources and References field so the provenance trail remains visible.
Rather than relying on a fixed query list, the search process widens iteratively to exhaust coverage within the tracked domain. Each run records every new reference as a lead and resolves it in the current run or a subsequent run, continuing until the lead pool is drained.
High-priority sources — statements by senior officials, major regulatory frameworks, primary regulatory bodies — remain in the active search queue even when a run returns no results. Subsequent attempts use alternate name forms, broader keyword variants, venue cross-products, and direct parses of official institutional channels, as these sources are deemed important enough to warrant a more rigorous search. This priority list is maintained and expanded as new high-signal sources are identified.
High-signal sources consist of new or amended regulatory frameworks, statements by tracked actors expressing directional views, analytical reports on governance dynamics, international position papers, and substantive policy commentary that reveals how relevant experts frame AI issues.
Source selection applies a relevance gate on policy signal, testing whether a given source reveals something meaningful about how the jurisdiction is thinking about, regulating, deploying, or competing on AI. Through our tag system, explained below, we outline the key concepts within AI policy that users may want to filter for. By defining what constitutes a meaningful policy signal, we can effectively filter sources on whether they answer the following question: would a researcher tracking AI governance find this informative? Documents carrying no policy signal are excluded to reduce noise. Excluded categories include routine filing batch lists, compliance how-tos, individual company filing announcements, and news rewrites that restate a regulation without original analysis. Borderline cases are marked for manual evaluation.
Source Types
Tag Design Principles
Tags are organized across three categories: Governance Domains, International Dimensions, and Risk Framing. A single entry can carry multiple tags within a category. The taxonomy is designed under a few explicit principles with consideration for the challenges of cross-lingual and cross-contextual coding, in an attempt to be as context-agnostic as possible.
For Governance Domains and International Dimensions, tags are applied only when the relevant domain is explicitly addressed in the document, not inferred from institutional affiliation or analytical judgment about what the document implies.
Risk Framing tags are implemented differently. These are applied when a document demonstrates awareness of a concern recognizable to the risk category, regardless of whether it is explicitly named.
To the best of our ability, each tag is designed with conceptual equivalence in mind — referencing the same phenomenon across cultural and jurisdictional contexts and avoiding Western-coded framing that would systematically miscategorize non-Western documents. For example, competition in Western policy discourse commonly denotes US-China strategic rivalry; in Chinese regulatory and industry discourse, 竞争 frequently appears in the context of domestic inter-firm competition between labs, a distinct phenomenon with different regulatory implications. The taxonomy therefore separates Great Power Competition — applied only when the document itself uses explicit rivalry framing (大国竞争、科技博弈、卡脖子) — from Industrial Policy, which captures domestic market development. Ethics & Values Alignment is similarly defined without presupposing any single normative tradition as a referent, to avoid treating Western AI ethics discourse as the conceptual default.
Governance Domains
International Dimensions
Risk Framing
Tag Distribution Across Sources
Bar length represents share of total entries carrying each tag. Entries can carry multiple tags, so totals across categories overlap.
Loading…
Acknowledgements
We are grateful to Parv Mahajan, whose early prototype and methodology informed the foundation of this project. We thank Edward and the Safe AI Forum for their feedback on source coverage and analytical framing. We also thank the researchers at RAND, CNAS, and the Institute for Progress who offered early input on the dashboard's direction.