How Can an Outsider Judge Whether a Car Without a Steering Wheel Is Safe?
By Scott Shields – Contributing Writer – Capitol Times Media - From Conversations and
Material of Zhu Weisha. Learn more about Zhu Weisha here at Capitol Times Media's July
Magazine Issue
Relying on Verifiable Thinking
A car with no steering wheel and no brake pedal pulls up in front of you.
The car company tells you it is safe. Engineers tell you it has undergone extensive testing.
Regulators have their own safety standards. The question is: you do not understand
autonomous driving, so why should you believe it is safe?
This was the real problem we encountered when we began studying autonomous driving.
We did not originally work in this field. Faced with the professional questions surrounding
Waymo and Tesla, we were outsiders.
We do not know how airplanes are built, yet we must decide whether they are safe to fly in;
we do not know how drugs are developed, yet we must decide whether to take them; we do
not know how large models are trained, yet we constantly hear claims that “artificial
intelligence (AI) may destroy humanity.” We cannot become experts in every field. Other
than trusting authorities and experts, do we have any other way?
That became the question that truly interested us after we began studying autonomous
driving.
1. Why Would a Car Without a Steering Wheel Be Unsafe?
When people first see an autonomous vehicle without a steering wheel, a natural reaction
is: if something goes wrong, a human cannot take over, so it must be more dangerous. That
sounds reasonable. But look carefully: it is no longer a fact. It is a judgment.
“There is no steering wheel” is a fact. Between that fact and “therefore it is unsafe” lies an
arrow.
So instead of continuing to argue about whether the steering wheel is important, we
changed the question: what safety function does the steering wheel serve? Once it is
removed, who or what takes over the safety functions previously performed by the driver?
For example: how does the car perceive its surroundings? What does it do when the
situation is uncertain? What happens when the system fails? What if the vehicle enters a
road environment it cannot handle? When must it slow down or stop? After an accident,
can we know what actually happened? At this point, “Is a car without a steering wheel
safe?” has changed from an intuitive question into a series of verifiable questions. We also
saw something clearly for the first time:
An outsider does not have to learn how to build an autonomous vehicle before
beginning to verify whether it is safe.
2. We Found That the Experts Had Thought Much More Deeply Than We Expected
As the research continued, the first thing we encountered was not a long list of expert
mistakes. Quite the opposite. Many of the questions we initially raised had already been
considered by the professional safety system.
Waymo uses the Operational Design Domain (ODD) to define the environments in which a
vehicle is qualified to operate. It uses a Safety Case to answer: why do we have reason to
believe the system is safe? What is the safety claim? What is the evidence? How does the
evidence support the claim? If conditions change, does the old evidence remain valid? If a
new accident occurs, should the previous safety judgment be reopened?
We then asked whether the Safety Case itself was reliable. We found that Waymo was
already studying how to assess the credibility of the Safety Case itself and had introduced
external audit.
At that point, our research discipline became:
If others have already done it, acknowledge that they have done it.
If the evidence is insufficient, write UNKNOWN.
If a judgment cannot survive a counterexample, change the judgment rather than the
evidence.
The deeper we went, the more clearly we saw the depth accumulated by a mature
professional system over more than a decade. But that did not end the questions.
3. Experts Know More Than We Do, but Their Conclusions Still Need Verification
This was an important distinction. In autonomous-driving expertise, we obviously cannot
compare with Waymo or Tesla engineers. But when a company moves from that
professional knowledge to the conclusion, “This system is safe enough to remove the
steering wheel and carry passengers on its own,”
that is no longer merely a professional fact. It is a judgment. And once it is a judgment, we
can keep asking:
What is the safety standard? What is the basis? What has testing covered, and what has it
not covered? Could different bodies of evidence share the same blind spot? How many
vehicles should be allowed onto public roads in the first exposure? What additional
evidence is needed to expand from small-scale operation to an entire city? When software
changes, how much of the old evidence can be inherited? How does the safety of one
vehicle become the safety of one hundred thousand vehicles running the same version?
What risks still cannot be calculated reliably?
These questions do not require the questioner to know how to design an autonomous
driving algorithm. They require another ability: to verify how a judgment was formed. We
gradually saw that professional knowledge and the final judgment are not the same thing.
Experts provide professional facts that we do not possess. But expert status itself is not
evidence.
4. How Two Outsiders Formed a Safety Judgment
There was another unusual feature of this research. It was not one outsider studying
autonomous driving. It was two outsiders: one human and one AI. Nor did we use the usual
pattern of “human asks, AI answers.” The actual process was quite different.
AI first expanded the field of material: Waymo, Tesla, the National Highway Traffic Safety
Administration (NHTSA), safety standards, real incidents, Safety Cases, regulation, and
adjacent theories. AI found angles the human had not considered. Those angles then
triggered new human questions.
At one point, AI identified what looked like an important risk. The human immediately
asked: Musk cannot possibly be operating without AI. Why would we be able to think of a
problem that they had not thought of? We cannot assume that an industry has missed a
problem merely because we have discovered it.
We searched and found that they had thought about it. We then attacked their handling
mechanism. If the mechanism held up, the issue should be removed from our “loophole
list.”
If the public evidence was insufficient, we left the result as UNKNOWN.
Only reasoning gaps that remained unclosed after those steps were entitled to remain.
So our research kept cycling:
AI expands human compresses, a candidate judgment forms, the strongest
counterexample is sought, professional evidence attacks our judgment, the judgment
is revised, verification begins again.
No participant in this process possesses automatic correctness.
Humans can form mistaken intuitions. AI can mistake correlation for causation. Experts
can also move from correct professional facts to a conclusion that has not yet been
sufficiently demonstrated. What matters, therefore, is not “who is smarter,” but whether
the collaborative mechanism can keep discovering its own errors.
5. What Can We Know in the End?
By this point, we had not become autonomous-driving experts. But we could form a much
more precise judgment than the intuition with which we began.
The absence of a steering wheel does not, by itself, prove that a car is unsafe.
A steering wheel is only one way in which a human assumes control responsibility in a
conventional car. If the functions previously performed by the driver, perception, judgment,
control, fault handling, and risk exit, can be reliably replaced by other mechanisms, and
those mechanisms have been sufficiently verified within explicit operating boundaries,
system versions, and safety regimes, then a car without a steering wheel can reach:
a level of safety sufficient to permit operation.
This does not mean “autonomous driving is absolutely safe.” Nor does it mean “the
company says it is safe, therefore it is safe.”
What we can form is a bounded judgment: which safety functions have relatively strong
evidence; which problems already have mature engineering solutions; which conclusions
hold only within a specific Operational Design Domain (ODD), software version, and
operating regime; which risks still lack sufficient public evidence; and which questions can
presently only be marked UNKNOWN.
This is what we mean by verification. Verification does not guarantee that every question
will end in affirmation or rejection. A judgment may receive three results:
Affirmed. Rejected. Uncertain. Knowing what we do not know is itself a result of
verification.
6. Autonomous Driving Was Not the Only Thing Being Tested
At this point, we realized that two things had actually been under test.
The first was autonomous-driving safety. The second was our own method of knowing.
Neither of us originally understood autonomous driving. We did not hand over the right to
judge simply because we were outsiders; nor did we pretend to be autonomous-driving
experts simply because we had a method. We did something else:
We began by seeking professional knowledge, let AI expand the radius of cognition, let
the human repeatedly compress and form judgments, and then let evidence and
counterexamples attack those judgments. Facts, judgments, conclusions,
overturning, and revision kept cycling.
What remained at the end was not “whom do we believe?” It was what we have reason to
believe, why we believe it, and under what conditions that judgment could fail. The
significance of this may extend far beyond the safety of autonomous vehicles.
Ordinary people encounter professional judgments they do not understand every day. In
the past, we easily reduced the difficulty to one question: which expert should I trust? Our
autonomous-driving research suggested another possibility. Perhaps the question should
not be: who is qualified to tell me the answer?
It should be: what process must this answer pass through before I have reason to accept
it?
7. This Was Also an Experiment in Cognition
AI has sharply reduced the cost for an ordinary person to enter an unfamiliar professional
field. Large bodies of material that once took a long time to reach can now be expanded
quickly. But easier access to information does not automatically produce reliable
understanding. Quite the opposite. An AI with an enormous associative radius can rapidly
generate many explanations that sound plausible. Without verification, it can also
manufacture errors more efficiently. The important thing is therefore neither to let AI judge
in place of humans nor to reject AI, but to build a human-AI cognitive mechanism that can
keep correcting itself.
AI can provide angles the human had not considered. The human can judge which
possibilities are worth pursuing. AI searches for evidence and counterexamples. The
human forms a new judgment. New facts then return to attack both sides. Cognition is no
longer simply:
Question — Answer
It becomes a loop that can be reopened again and again. We originally used this method to
verify autonomous driving.
In the end, autonomous driving also verified the method in return. An autonomous vehicle
faces a world that can never be fully predicted, yet every second it must decide what to do
next. A person entering an unfamiliar field faces a similar problem. We can never possess
all knowledge. Yet we still have to form judgments.
For the machine, the question is: how does it earn the authority to act in an uncertain
world?
For the human, the question is: how do we earn the authority to judge in a world we do not
fully understand?
That may be the most important question this study of autonomous driving has left us.
Our article “Lessons from Autonomous Vehicles: How to Build Verifiable Safety in an
Uncertain World” was the answer produced by two outsiders. The Verifiable Thinking we
used can currently be compressed into two high-level rules, together with one special case
raised by autonomous driving:
For determinate questions, verify what is true and what is false.
For uncertain questions, verify the basis of the judgment and its probability of error.
Autonomous driving presents a special case: the probability of error itself cannot be
estimated reliably, yet reality still requires an action. In that situation, verifying a
single judgment is not enough; the regime that produces the action must also be
verified.
We are not yet concluding that this is an independent third rule. It may instead be a
further extension of the second rule under conditions in which action cannot be
deferred.
For more than a decade, frontline autonomous-driving experts have been solving safety
problems in the real world. Many of the problems they have addressed through engineering
practice converge with the questions we reached from Verifiable Thinking, though by a
different route.
Recently, we have used more than ten articles to examine influential industries and figures
as tests of Verifiable Thinking. Autonomous driving is the field in which we found the fewest
faults; after verification, our assessment of their work rose rather than fell. Respect to
Waymo. Respect to Tesla.



