Who Rates the Rating Agencies?
- Capitol Times Media

- 11 minutes ago
- 14 min read
By Scott Shields – Contributing Writer – Capitol Times Media - From Conversations and Material of Zhu Weisha. Learn more about Zhu Weisha here at Capitol Times Media's July Magazine Issue
A Testable Baseline for the Warsh Fed’s Reaction Function
From America’s AA+ to a Verifiable Credit Market in the Age of AI
In 2011, Standard & Poor’s removed the United States from the AAA sovereign rating it had long held. In 2023, Fitch also downgraded the United States from AAA to AA+. In 2025, Moody’s lowered the U.S. long-term issuer and senior unsecured ratings from Aaa to Aa1.
Yet U.S. Treasury securities did not cease to be core global safe assets. They remain among the world’s most important reserve assets, liquid assets, and forms of collateral. This raises an interesting question: if Germany is AAA and the United States is AA+, does that mean German government debt is safer than U.S. Treasury debt in every relevant sense?
Of course not.
The issue is not simply that the rating agencies “got America wrong.” Rating agencies have never claimed that a single AAA or AA+ can answer every question about a country, a company, or a bond. A credit rating mainly assesses the relative credit risk that an issuer or obligation will fail to meet its financial obligations in full and on time. That is one dimension of the broader credit problem; a rating agency provides one important measuring stick among several.
The Congressional Budget Office (CBO), the Government Accountability Office (GAO), the Federal Reserve (Fed), the U.S. Department of the Treasury, rating agencies, and market investors all study the United States, but they are not asking exactly the same question. Long-term fiscal sustainability is one question. Whether a very large stock of debt can be continuously refinanced is another. Whether the U.S. Treasury securities market can continue to function under stress is yet another. And whether the U.S. government will meet its debt obligations on time is still another. Credit is multidimensional; different dimensions require different measures.
This also explains how the United States can lose AAA status and still retain one of the world’s most important safe assets. Markets may criticize rating agencies, but that does not make ratings meaningless. The more common mistake is ours: treating the agency’s measuring stick as though it represented the whole picture.
The question therefore moves one step further. If real-world credit is inherently multidimensional, why do we ultimately compress such a complex world into a handful of symbols—AAA, AA+, AA, and so on?
1. AAA Is Not a Mistake; It Is a Way Humans Manage Complexity
Knowledge and research relevant to credit are dispersed across different institutions.
Rating agencies study repayment capacity. Fiscal authorities study debt. Central banks study financial stability and market liquidity. Investors study prices. Scholars study safe assets and capital flows. The same is true for a company: finance teams know its balance sheet, banks know its credit lines, suppliers know its payment behavior, customers know its delivery performance, auditors know its financial statements, and rating analysts study its industry, competitive position, and capital structure.
An ordinary investor cannot continuously track dozens of institutions, hundreds of indicators, and constantly changing assumptions. Even within professional institutions, complex judgments must be converted into results that can be communicated, cited, and acted upon.
Human beings therefore developed an important engineering solution: compression.
AAA, AA+, AA, and A are a compressed language.
That compression is not obsolete. Contracts need clear standards. Investment mandates need boundaries. Index construction, collateral rules, risk limits, and internal governance also require a common language. No matter how capable AI becomes, an investment committee cannot debate hundreds of credit nodes every time it buys a bond.
We will still need compression to reach a concise conclusion. But with AI, the complex world that has been compressed no longer has to disappear; it can be reopened when needed. AAA will not vanish. AA+ can remain, while the facts, assumptions, models, judgments, dissenting views, and historical changes that produced AA+ are preserved underneath. A user who wants a quick conclusion can stop at AA+; a user who wants to know why can keep going.
This may look like a small change, but it changes the structure of the rating product.
2. A Rating Can Become an Entry Point That Can Be Continuously Expanded
If AI merely recalculates an “overall U.S. credit score of 94,” it has not achieved very much. It has only made the old ruler more complicated.
The more valuable approach is to make the same AA+ expandable.
That expansion has at least three meanings.
The first is to expand a single letter grade into the different credit states that sit behind it.
The same United States can simultaneously have a worsening long-term fiscal position, rising rollover pressure, strong productive capacity, strong Treasury financing capacity, extremely deep liquidity in normal markets, pockets of fragility under stress, a currency at the center of the global monetary system, and recurring political risk around the debt ceiling.
All of those judgments can be true at the same time. There is no need to force them into a single “complete answer” such as 92.7 points.
Traditional ratings can continue to perform the compression function, while a state layer preserves the underlying complexity. The two are not in conflict.
The second meaning is that the same factual base can be expanded differently depending on the question.
An investor preparing to hold a 30-year U.S. Treasury bond may ask, “What is my biggest risk?” The system would naturally bring long-term fiscal policy, real interest rates, inflation, productivity, the long-run tax base, and the dollar system to the foreground.
Another investor may hold a three-month Treasury bill (T-bill) and be asking about near-term debt-ceiling risk. In that case, the relevant variables may be the specific maturity date, the Treasury’s cash balance, the “X-date” on which the Treasury estimates that available cash and extraordinary measures could be exhausted, payment arrangements, and short-term market liquidity.
If the user is a central bank reserve manager, the main questions may instead concern market capacity, liquidity, the dollar payment network, reserve-asset functionality, and the collateral system.
The underlying country is the same and the factual base is shared, but what needs to be expanded is different.
This is one of AI’s biggest advantages over a standard PDF report. It is not a license to create different facts for different users, nor to quietly alter the methodology behind the formal rating. It allows the system to retrieve from the same factual base the parts that are most relevant to the question at hand.
The third meaning is that the judgment itself can be expanded through time.
When a user sees AA+ today, the questions should not stop at “Why AA+?” The user should also be able to ask: What are the three most important judgments behind it? Which judgment is most uncertain? What new fact would make the agency acknowledge that its current view needs to change? Why did it not hold the same view last year? Which forecast from last year failed to materialize?
At that point, what is being expanded is no longer just data. It is the rating agency’s own judgment history.
That layer matters greatly.
In credit evaluation, the hardest thing to verify is usually not whether a number was copied correctly. It is why a judgment was formed at the time and why it later changed.
3. AI Alone Is Not Enough: What Is Expanded Must Also Be Verifiable
Large models can read enormous amounts of material and rapidly organize it into a logically coherent story.
But a coherent story is not the same as a reliable judgment. AI can forget what it previously believed, smuggle later facts into an earlier forecast, blur the distinction between facts and assumptions, or apply different standards at different times. Several AI systems giving the same answer do not automatically make that answer more reliable: they may be using the same data, the same article, or simply repeating the same evidence chain. A real AI credit system therefore cannot preserve only today’s answer. It must also know where the original facts came from, what data definitions were used at the time, which inputs were facts and which were assumptions, what model was used, what the analyst concluded, and why the committee accepted or modified that conclusion. If the model later changes, the old version should remain replayable. If the judgment changes, the system should record which new facts drove the change.
If the final judgment proves wrong, we should at least be able to ask again: Was the error in the facts, the model, the assumptions, or the human judgment?
That is the dividing line between a verifiable system and ordinary “AI ratings.”
Deterministic facts can be checked, and processes can be replayed. Cognitive judgments cannot be proved “true” in the same way as a bank balance, but the evidence, process, and history by which those judgments were formed can be reviewed.
Credit ratings are a classic case of cognitive judgment. A verifiable rating system therefore does not claim, “I can prove AA+ is absolutely correct.” It allows later users to see what the agency relied on when it assigned AA+, what conditions supported that view, what conditions would weaken it, and what new facts should trigger a change.
4. The System Does Not Need to Start Big
If a rating agency wanted to experiment with this approach today, I would not begin by building a grand “global credit operating system.”
The first step can be small.
The traditional “AA+ / Stable” can remain. It would already be valuable if the user could see four additional things behind it: the key assumptions on which the rating rests, the biggest uncertainties today, the facts that would invalidate the current judgment, and the history of how the agency has revised its view.
After seeing AA+, the user can ask, “Why?”
The system expands the most important facts and judgments.
The user can then ask, “Which part is most uncertain?”
The system identifies the weakest part of the current judgment.
Then: “Under what conditions would you change your view?”
The system retrieves the weakening or invalidation conditions that were recorded in advance.
And then: “Why did you not think this way last year?”
The system retrieves the earlier version and the reasons for revision.
Even this would be more valuable than inventing a new “U.S. credit score of 94.”
It would not overthrow the existing rating system. It would open the judgment process that is currently compressed beneath a letter grade.
Cross-validation should follow the same principle. The fact that three AI systems all say AA+ is not important by itself. First, we should check whether the underlying facts come from relatively independent sources. Next, we should compare how different models interpret the same facts. Then we should see whether the rating judgment shows an abnormal divergence from market signals such as yields, credit default swaps (CDS), the repo market, and Treasury auctions, divergences that deserve explanation. Only after that should different AI systems be asked to interpret the materials.
Market prices are not the “ground truth” of a rating, and data from different government agencies are not statistically independent merely because they come from different institutions. What matters is how independent the sources, methods, and objects of observation really are, and where common dependencies remain.
At times, disagreement is more valuable than agreement. A conflict is often the best entry point for the next round of research.
5. Verifiability Must First Constrain the Rating Agencies Themselves
The rating industry has another issue that cannot be avoided: rating agencies do not stand outside all conflicts of interest.
In the United States, credit rating agencies operate within the Nationally Recognized Statistical Rating Organization (NRSRO) registration and oversight framework, and ratings are widely used in investment policies, financing arrangements, risk management, and other institutional settings. At the same time, much of the industry has long relied on the issuer-pay model, under which issuers, underwriters, or related debtors pay for the rating. Regulators have long treated the potential conflicts created by this business model as an issue that must be managed.
AI will not automatically remove those problems. In fact, “conversational ratings” may create new ones.
If an issuer can keep testing assumptions with a rating AI until it finds the most favorable explanation; if a paying client can receive material credit judgments that others cannot see; or if the commercial side of the firm can influence which facts the AI emphasizes and which risks it plays down, greater technological sophistication may simply make the problem harder to detect. The system therefore needs clear boundaries.
A common factual layer, the formal rating, the public explanation, and the client’s own scenario analysis should not be mixed together. A client can certainly ask, “What happens if revenue falls 20% next year?” But that scenario analysis must not quietly turn into a different formal rating available only to that client.
A rating committee can decide to override a model recommendation, but the reason should be recorded.
AI can participate in judgment, but it cannot become a black hole of responsibility.
Professional analysts, rating committees, compliance functions, and regulators do not lose their value because AI appears. AI is better used to preserve the professional process that is often hidden beneath the final letter grade, so that authorized reviewers can re-enter that process when necessary.
This answers half of the question in the title: before rating others, rating agencies should bring their own judgment processes into a verifiable era.
6. The New Market Rating Agencies May Really Open: Corporate Credit
If the article ended here, it would still be only an article about improving rating products.
Once a corporate model enters the framework, however, the boundaries of the market begin to change.
Existing rating systems already distinguish between issuer credit and the credit of a specific debt instrument. Assessing a company’s overall capacity to meet financial obligations is not exactly the same as assessing one particular bond it issues. A specific bond must also be analyzed in terms of payment priority, collateral, guarantees, and contractual structure.
Corporate credit is therefore already an important foundation for bond ratings.
That leads to a natural question: if a company already has a continuously maintained AI credit model, why should every new bond issuance require analysts to relearn the company from scratch?
A company’s balance sheet, cash flow, debt maturities, industry conditions, competitive position, management, major litigation, and guarantees are continuously changing. If those facts are maintained in a persistent corporate credit base, a new bond can be analyzed by adding the structure of that particular instrument to the underlying issuer view.
Bond ratings could gradually become one output of a broader corporate credit model.
In some respects, companies are better suited than sovereigns to this approach. A sovereign involves taxing power, monetary capacity, political institutions, policy willingness, social stability, and geopolitics—many of which have no uniform accounting framework. A company at least has a defined legal entity, a balance sheet, an income statement, a cash-flow statement, audited accounts, contracts, and relatively clear boundaries of responsibility.
Of course, feeding several financial statements into an AI does not automatically produce a reliable AAA.
But much of the repetitive groundwork that analysts once had to perform manually is increasingly machine-processable. As long-run financial data, debt changes, capital expenditure, major transactions, litigation, guarantees, industry conditions, and competitive relationships are continuously added, a persistent corporate credit state can begin to emerge.
Nor should an “AAA company” be understood to mean “the best company” or “the best stock to buy.”
The term refers only to credit. A slow-growing, modestly valued company with little debt and stable cash flow may have excellent credit, while a fast-growing, technologically leading company may still fall short of the highest credit grade because of its capital structure or cash-flow risk.
Equity investment value and debt-repayment credit remain two different measuring sticks.
The market really opens with the next question: if a corporate credit model can persist over time, why should only companies preparing to issue bonds be worth evaluating?
A company that has never issued public debt still needs bank loans. Suppliers must decide whether to extend trade credit. Customers may need to decide whether to prepay. Insurers must decide whether to underwrite. Other companies must decide whether to sign long-term contracts. Government procurement teams must judge whether a contractor will still be able to perform years from now.
For example, if a supplier is deciding whether to grant a customer 90-day payment terms, the system can bring short-term cash flow, near-term debt maturities, bank credit lines, accounts payable, and payment history to the foreground, while keeping long-horizon equity valuation out of the center of the analysis. The corporate credit base has not changed; only the parts most relevant to this particular credit relationship are being called.
These economic activities involve credit judgments every day.
Today, different institutions repeatedly collect the same information and build their own risk controls. One reason is that human evaluation is expensive: there has been no sufficiently low-cost common credit base that can be continuously maintained and called across different contexts.
AI lowers the cost of evaluation; verifiability gives machine-generated evaluation an evidentiary basis and a history.
Together, they could gradually shift the size of the rating market away from “how many bonds need ratings” toward “how many economic relationships require a credible judgment of a counterparty.”
Historically, financing demand created the rating market. In the future, demand for verification may create a new market as well.
7. One Step Further: The Object of Evaluation Can Go Beyond Solvency
Corporate ratings can be extended further, but there is no need to finish that discussion in this article.
Whether an asset manager can pay its own debts is a credit question. But investors also care about whether it manages client assets under the rules it promised to follow, whether authorization procedures are actually enforced, whether risk rules exist only on paper, and whether failures leave a process that can be replayed. That is no longer purely traditional debt credit; it is closer to the operating or institutional credibility of an organization.
The same logic can extend to people.
There is no need to create a single “credit score of 72” for a person. Anyone who makes public judgments over a long period—an entrepreneur, fund manager, economist, analyst, or political figure—leaves a large public record. Statements of fact, forecasts, promises, later outcomes, and subsequent revisions are different categories and should remain different.
AI makes it possible to organize those records over long periods at far lower cost.
There is an old Chinese saying: “Heaven keeps the record.”
Historically, that was mostly a moral expression, because human memory is limited and, with enough time, many old statements simply disappear from attention.
With AI, long-term preservation and retrieval of public records become much easier. AI is not Heaven, of course, and it can itself remember incorrectly. Any important record should therefore remain traceable to the original statement, original document, and original point in time.
Once people become accustomed to asking a rating agency, “Why did you think this five years ago?”, it is natural to ask why the same question cannot be posed to a company, an institution, or anyone who makes public judgments over time.
Rating is simply the most suitable place to begin.
8. Who Rates the Rating Agencies?
We can now return to the original question.
Rating agencies rate countries, companies, and bonds. Who, then, rates the rating agencies?
If the answer is simply “another, more authoritative institution,” we enter an endless regress.
If the answer is AI, we have only replaced trust in an institution with trust in a machine.
A more interesting answer may lie in the rating agencies’ own histories. What facts did an agency have ten years ago? What judgment did it make at the time? Which judgments later proved useful? Which were wrong? Why? When did the agency realize it was wrong? Did it change its model? Did it improve the next time it faced a similar situation?
When those records persist over time, what a rating agency has accumulated across decades is no longer just a succession of AAA and AA+ grades. It is a highly valuable history of “facts—judgments—outcomes.”
That history may become one of the most important assets of a rating institution in the AI era.
A general-purpose large model can certainly read today’s financial statements. But a rating agency’s deeper advantage may be that it possesses a professional record of “this is what we knew then, and this is how the world later turned out.” If that history is structured, while preserving the version that existed at the time rather than rewriting it in light of later outcomes, it becomes extraordinarily valuable material for training, calibration, and verification.
The core capability of rating agencies in the past was to compress a complex world into a credible letter grade.
The best rating institutions of the future may add another capability: continuously preserve the world behind the letter grade, reopen it when needed, and allow their judgments to be tested by subsequent facts.
As corporate credit models expand, bonds may become only one output. Banks, insurers, supply chains, investors, procurement teams, and other institutions may all call a company’s credit state. If AI agents become deeply involved in economic activity, the entity calling a credit judgment may not even be a human. A procurement agent could check a counterparty before signing a contract; a banking agent before lending; an insurance agent before underwriting.
At that point, a rating may no longer be only a report. It may gradually become a form of credit infrastructure embedded in economic activity.
AI makes it possible to preserve and retrieve a complex world at scale; verifiability constrains that capability. Whether a rating institution can truly organize AI, historical data, professional judgment, and verifiable mechanisms will help determine whether it can evolve from a rating service provider into a broader credit infrastructure.
Whether the major rating agencies continue to deserve market trust should not depend only on the ratings they assign to others today. It should also depend on whether their own judgments withstand long-term verification. If an institution’s accumulated record and verifiable track record support it, we have good reason to continue trusting it.


