There is an emerging network of independent AI auditing organizations. And as AI diffuses into more sectors of the economy, it's likely that more existing organizations (firms across all sorts of domains, public bodies, etc.) will also need to develop "AI evaluation skills."
Thus, we are going to develop some globally distributed capacity for AI evaluation (and this is good).
As policymakers begin to explore a design space of "AI-targeted policy" (e.g., AI taxes and dividends, etc. -- which may not be necessary in all cases, as tried-and-true economic policy may work well for addressing certain issues), they are going to face various scoping questions. Which firms count as AI firms (e.g., who is going to face an "AI tax" if such a thing were to be implemented)? Which models count as frontier models?
In the past, we've seen big debates around ideas like FLOP cut-offs.
The goal of this post is to make a pretty short, and I think uncontroversial argument: We should utilize this network of organizations with an interest and incentives to do independent evaluation to perform capability measurement that can be used to scope which AI developers would be affected by any kind of "AI economic impact policy" that targets "AI companies" or "AI labs".