top of page

Not Another Yardstick for the Rating Agencies

11 minutes ago
9 min read

By Scott Shields – Contributing Writer – Capitol Times Media - From Conversations and Material of Zhu Weisha. Learn more about Zhu Weisha here at Capitol Times Media's July Magazine Issue


A Review of Who Rates the Rating Agencies?


Who Rates the Rating Agencies? is, on the surface, about credit ratings. Look more closely, however, and it reaches a larger question: in the age of AI, what gives us reason to trust an institution, a company, or even a person who makes public judgments over a long period of time?

The essay begins with U.S. sovereign ratings. The United States has lost the highest rating from all three major agencies, yet U.S. Treasuries have not lost their place at the core of the world’s safe-asset system. This fact is often used to argue that rating agencies are “useless” or that


they “got America wrong.” This essay does neither.

It begins by acknowledging that ratings are useful.


AAA and AA+ are not, and were never meant to be, complete verdicts on every aspect of a country’s credit condition. Whether fiscal policy is sustainable, whether debt can continue to be rolled over, whether the Treasury market remains liquid, and whether the government will pay on time are related but distinct questions. A rating agency uses one important yardstick among several.


That starting point may look ordinary, but it matters. Only by admitting that ratings serve a real purpose can reform avoid becoming an exercise in tearing down an old system merely to replace it with a new one.


The essay’s real originality begins with the idea of “compression.”


AAA is a form of compression. Reality is too complex, so markets reduce large amounts of information to a few letters that can be used in contracts, investment mandates, and risk management. Until now there was little alternative, because no human being could continuously track hundreds of changing variables.


AI changes this for the first time. It does not mean we stop compressing. We can still assign AA+.

But once a complex reality was compressed into AA+, much of what lay behind it became hard to see: why AA+, where the strengths and weaknesses were, which elements were assumptions, which judgments were most uncertain, and why the conclusion differed from

last year’s. Those things were scattered across reports, models, committee records, and analysts’ minds.


Now they can remain behind the rating.


So the article is not really proposing that “AI can help rating agencies calculate more accurately.” It proposes a different structure:


Ratings can still compress reality, but compression no longer has to mean permanently losing the complexity underneath.


I think this is the first idea in the essay that deserves to be remembered.


The next step is “expansion.”


Expansion here does not mean attaching dozens of tables to AA+. It means three different things.


First, one letter can be expanded back into several conditions that coexist. U.S. fiscal conditions can deteriorate without eliminating its underlying economic strength; debt pressure can rise while the Treasury market remains extraordinarily liquid; political risk can exist while the dollar remains the core currency of the global system.


All of these can be true at the same time. Why force them into a single average?


The second kind of expansion is even more useful.


Different people ask different questions about the same United States.


An investor holding thirty-year Treasuries cares about long-term fiscal conditions, inflation, real interest rates, and productivity.


Someone holding a three-month Treasury bill and facing a debt-ceiling episode cares much more about the maturity date, the Treasury’s cash balance, and the X-date.


A central bank managing reserves looks at still another set of issues: market depth, the dollar system, and collateral functionality.


The facts have not changed. The question has changed, so the relevant facts change with it.


A traditional rating report had to be written once for everyone. AI can do something different: maintain a common factual base and let different users interrogate it according to the question they actually face. That is no longer simply “a better report.” It is a different product architecture.

The third kind of expansion is time.


This is the one I find most interesting.


When we see AA+ today, we need not stop at asking “Why AA+?” We can also ask:


Why did you not make the same judgment last year? What assumptions mattered most at the time? Which forecasts failed to materialize? When did you change your view, and why? Once those questions can be answered, a rating agency no longer leaves behind only a sequence of


AAA and AA+. It leaves behind a history of judgment.


This is where the essay truly reaches the idea of verifiability.


When people talk about AI and verification, they often slip into a basic error: assuming that enough data and enough modeling can prove that a judgment is true. Credit ratings plainly cannot work that way.


The future has not happened yet. How could anyone prove in advance that AA+ is absolutely correct?


The essay draws this boundary clearly.


Facts can be checked and processes can be replayed, but a judgment cannot be proven true in the same way as a bank balance. What can be examined is what was known at the time, what assumptions were used, how the judgment was formed, and what conditions should have caused it to change.


In other words, verifiability does not guarantee that you will always be right. It requires that you cannot make a judgment and then hide how you got there. You can say AA+ today, but you must leave the reasoning behind it.


When facts later change, others can return and ask why you said what you said at the time.

I think this matters far more than “AI improving rating efficiency,” because it touches the nature of judgment itself.


The original essay is also practical in one important respect. It does not leave verifiability at the level of principle. It reduces the first workable step to four things: key assumptions, uncertainties, invalidation conditions, and revision history. If a rating can expose those four elements, verifiability has already moved from an abstract principle to a product that can actually be built.


Another line in the essay is worth keeping: agreement among several AI systems does not make a conclusion more reliable by itself.


Many people still do not fully understand this.


If three models all say AA+, that sounds like cross-validation. But if all three rely on the same data, the same article, and similar training material, then we do not really have three pieces of evidence. We have one piece of evidence repeated three times. The real question is whether the evidence is independent, whether the models are meaningfully different, and whether the market is showing anomalies that deserve explanation, not how many AIs agree.


Sometimes disagreement is more useful than agreement, because disagreement tells you where the problem may be.


This suggests a research habit worth preserving: before counting how many answers agree, first ask whether the evidence is genuinely independent.


The essay becomes even more interesting when it refuses to apply verifiability only to the entities being rated. The rating agencies themselves must also enter the frame. That brings us back to the title.


There is another part of the original essay that should not be skipped: rating agencies have interests of their own. The issuer-pay model already creates potential conflicts. Add AI, and if issuers can repeatedly test assumptions or paying clients can receive different interpretations, “smart ratings” may simply make rating shopping harder to see. Common facts, formal ratings, public explanations, and client scenario analysis therefore need clear separation. If a rating committee overrides a model, the reason should remain on the record. Verifiability cannot demand transparency only from the rated party; the rater must first leave its own chain of responsibility.


Who rates the rating agencies?


If the answer is to find another, more authoritative rating agency, the chain never ends. If the answer is AI, we have merely exchanged human authority for machine authority.


The essay’s answer is: look at the agency’s past judgments. What did it know ten years ago?


What did it conclude? What happened afterward? Where was it wrong? When did it acknowledge the mistake? Did it improve? This is more interesting than asking “Who regulates the rating agencies?” because it does not search for a higher judge. It turns the agency’s own history into evidence.


There is also a significant industry insight hidden here. People often assume that AI will reduce the value of traditional rating agencies. Perhaps not. What may be most valuable is not today’s AAA or AA+, but decades of records showing “what we knew then, how we judged it, and what happened afterward.”


A general-purpose large language model can read today’s financial statements easily enough.

But it does not naturally possess an analyst’s contemporaneous judgment from decades ago. If those records can be organized, preserved in their original versions, and protected from hindsight rewriting, they may become some of the scarcest data in the AI era. Old rating-agency data may not be a burden at all; it may become more valuable.


The essay then makes another large move.


Why should the rating market revolve mainly around bonds? Even a company that never issues public debt is judged for creditworthiness every day. Should a bank lend to it? Should a supplier offer ninety-day terms? Should a customer prepay? Should an insurer underwrite it? Should a government award it a long-term contract? All of these are credit decisions.


If AI lowers the cost of maintaining a company’s credit condition continuously, enterprise credit could become a persistent base layer.


When the company issues debt, the bond rating would simply be one output of that base.

A supplier considering ninety-day terms would call up short-term cash flow, near-term maturities, bank facilities, and payment history.


A bank considering a five-year loan would call up a different set of factors.


The company is the same. The factual base is the same. What changes is the question.

At this point the essay makes one of its most imaginative claims:


In the past, rating markets were created mainly by financing demand. In the future, verification demand may create an additional market.


I think this is the essay’s most important commercial proposition.


If it proves true, the rating industry will no longer ask only, “How many bonds need ratings?” It will begin asking, “How many transactions need a trustworthy judgment?”


The economics of the market would change. But this is also the essay’s biggest unresolved question.


Banks, suppliers, insurers, and government procurement departments may all care about the same company, but they bear different risks, operate on different time horizons, and face different legal responsibilities.


So I do not expect a universal rating to make decisions for everyone.


A more plausible future is that institutions share a verifiable base of enterprise facts and credit conditions, while each makes its own final judgment.


The base can be common. The conclusion cannot.


That distinction will matter greatly if someone actually turns this idea into a product.

The essay briefly extends the idea to institutions and people, and I think it stops at the right point.


There is no need to assign a person a “credit score of 72.”


What matters is whether we can preserve, over time, what someone said, predicted, promised, what later happened, and whether they corrected themselves when they were wrong.

The phrase “Heaven keeps the record” sounds almost playful here, but it captures a very real problem.


Society used to forget.


People said things, and years later no one looked them up.


Predictions failed, and later people remembered only the successful ones.

An institution changed management, and past mistakes could easily be repackaged.

AI may first change not the quality of judgment, but the cost of memory.

But remembering is not the same as remembering correctly. That is why verifiability is necessary. AI can retrieve and organize.


The original facts, words, dates, and records must establish whether the memory is accurate.

These two capabilities together are what the essay is really trying to get at.

So if I were to classify Who Rates the Rating Agencies?, I would not call it an ordinary essay on rating reform.


It really contains three new ideas.


The first is:

AI can separate compression from forgetting for the first time.


The second is:


Once a judgment has a complete history, later facts can continuously calibrate it.

And the third is:


As the cost of verification falls, verification itself may generate new credit demand.

The first two change the product and the method.


The third changes the market.


Beneath all of this lies another question. Why do we trust an institution? Because of licenses, brands, experts, and long-standing reputation. Those sources of trust will remain.


But the essay adds another reason that was previously difficult to make available at scale:

I can now go back and inspect how you judged, what evidence you used, and what happened later. If that record survives long-term scrutiny, I have one more reason to trust you. That is where the essay reaches beyond the rating industry.


It does not say that verification replaces trust. It says:


Credit can arise not only from identity and authority, but also from the sustained verification of long-term, verifiable records.


If that proposition is ultimately borne out in practice, Who Rates the Rating Agencies? will leave behind more than an AI product concept for rating agencies.


It will leave us with a larger question:


When we decide whether to trust a decision-maker, can long-term verifiable records become one more source of evidence for that trust?

READ NEXT

Heading 2

Disclaimer:
 

The views and opinions expressed in the articles or Interviews published in this magazine are solely those of the respective authors and do not necessarily reflect the official policy or position of the Capitol Times magazine or Capitol Times Media , its editors, or its staff. The authors are solely responsible for the content of their articles. The magazine strives to provide a platform for diverse voices and opinions, and we value the principle of free expression. The magazine assumes no responsibility or liability for any errors or omissions in the content of the articles. In no event shall the Capitol Times magazine or Capitol Times Media be liable for any special, direct, indirect, or incidental damages. Furthermore, the inclusion of advertisements or sponsored content in Capitol Times magazine does not constitute an endorsement or guarantee of the products, services, or views promoted by the advertisers. Readers are encouraged to conduct their own research and exercise caution when making decisions based on advertisements or sponsored content featured in this publication.

Thank you for reading and engaging with our publication. Your feedback is valuable to us as we continue to provide a platform for thought-provoking content and diverse perspectives.

 

Capitol Times Media is a privately owned and independently operated media that publish Capitol Times Magazine. It is not affiliated with, endorsed by, or connected to the United States government, the U.S. Capitol, Congress, or any federal, state, or local government agency. Content published by Capitol Times Magazine includes both editorial content and sponsored or paid content.


© 2026 by Capitol Times Media LLC - Privacy Policy

bottom of page