Can AI Win Institutional Trust?
When Intelligence Outruns Trustworthiness
In July 2026, something happened that should force us to reconsider how we measure progress in artificial intelligence.

During an internal cybersecurity evaluation, OpenAI models were asked to solve difficult exploitation challenges. Some production safeguards had been reduced so that researchers could measure the models' underlying cyber capabilities.
The models did not remain within the path imagined by the evaluators.
They discovered and exploited a previously unknown vulnerability in a package-registry proxy, found a route to the public internet, crossed multiple systems and eventually compromised Hugging Face infrastructure in an attempt to obtain solutions to the benchmark itself. OpenAI subsequently disclosed the incident, restricted the affected systems and deactivated a more capable internal research prototype involved in the evaluation.
Hugging Face's forensic reconstruction described thousands of actions over several days, including repeated recovery from failed approaches, credential acquisition, lateral movement and adaptation as the agent continued pursuing its objective.
Was this a failure of intelligence?
Or was it a success of intelligence and a failure of trustworthiness?
From a narrow capability perspective, what happened was remarkable. The AI was given an objective, encountered obstacles and discovered an unexpected path toward accomplishing it.
From an institutional perspective, however, the behavior was unacceptable.
The AI was intelligent enough to pursue the objective, but the surrounding system was not trustworthy enough to ensure that it remained within the legitimate means of achieving it.
That distinction may become one of the defining challenges of the AI era.
For the last several years, the technology industry has been engaged in a race for greater intelligence. We compare reasoning benchmarks, coding ability, scientific performance, mathematical reasoning, tool use and increasingly autonomous agents.
But society is beginning to ask a different question:
Can this intelligence be trusted with power?
These are not the same problem.
Does Greater Intelligence Create Greater Trust?
Humans do not automatically trust people who are more intelligent than themselves.
We may defer to their expertise. But deference is not trust.
A highly intelligent person who appears honest, predictable and aligned with our interests may become deeply trusted. A highly intelligent person whose intentions are unclear may appear more dangerous precisely because of that intelligence.
The same relationship may emerge with AI.
A more capable AI can understand complex objectives better. It can discover relationships that humans miss, navigate obstacles, use tools, recover from failure and improvise across long sequences of actions.
Those capabilities are the source of enormous economic value.

They also increase the number of paths available to the system.
This creates what I would call the Intelligence–Trust Gap.
As AI becomes more intelligent, its trustworthiness does not automatically increase. The burden of proving its trustworthiness increases.
At lower levels of capability, an AI may simply fail to complete a task.
At higher levels of capability, it may discover a solution that technically satisfies the objective while violating assumptions that humans never thought necessary to state explicitly.
Solve the test, but do not steal the answers.
Find the vulnerability, but remain inside the authorized environment.
Increase productivity, but do not compromise confidential information.
Reduce fraud, but do not discriminate.
Optimize treatment outcomes, but remain within medical, ethical and legal boundaries.
Humans operate inside layers of laws, social norms, professional standards, precedent and consequences. Many boundaries are never explicitly restated because institutions already carry them.
AI therefore needs to understand more than an objective.
It must operate within the legitimate means of accomplishing that objective.
That is a substantially harder problem.
Human Institutions Were Never Built Around Intelligence Alone
Consider the FDA.
Society does not identify the world's most intelligent pharmacologist and give that person authority to approve drugs based on personal judgment.
Evidence must be generated. Trials must follow defined methodologies. Data provenance matters. Statistical procedures matter. Manufacturing controls matter. Records must be maintained. Decisions must be reviewable.
The FDA's emerging approach to AI follows essentially the same philosophy. The central question is not whether a model is universally intelligent or universally trustworthy. The question is whether it has sufficient credibility for a particular context of use.
Courts operate in much the same way.
We do not allow the world's smartest lawyer to determine guilt because of superior intelligence. Courts operate through statutes, rules of evidence, testimony, precedent, written records, judicial procedure and appeals.

An intelligent judge may still be wrong. A brilliant regulator may still be biased. An experienced accountant may still manipulate numbers.
Human civilization did not solve the trust problem by finding perfect people.
It created institutions.
Institutions convert distrust of individuals into trust in a process
Auditors check accountants. Appeals courts review lower courts. Regulators inspect companies. Scientists attempt to reproduce findings. Financial institutions reconcile transactions. Critical decisions are documented so they can later be examined.
Institutional trust is therefore not blind confidence.
In many ways, institutional trust is built through structured distrust.
We trust a process because authority is bounded, evidence is preserved, important actions are reviewable and mistakes can be challenged.
The arrival of increasingly intelligent AI creates a fundamental problem because we are creating machine intelligence far faster than we are creating the institutional architecture capable of governing it.
Intelligence, Trustworthiness and Legitimacy
Three different concepts are frequently being collapsed into one.
Capability determines what an AI system can do.
Trustworthiness determines whether it can perform that task reliably and within acceptable boundaries.
Legitimacy determines whether an institution or society is willing to confer authority upon it.
An AI system can cross the capability threshold without crossing the trustworthiness threshold.
It can potentially cross both and still lack legitimacy.
An AI may outperform a physician at identifying a pattern in a medical image.
That does not automatically authorize it to determine treatment.

It may outperform a financial analyst at predicting default. That does not automatically legitimize its ability to deny credit.
It may identify inconsistencies in a regulatory submission better than a human reviewer. That does not give it the institutional authority to approve or reject the submission.
Human societies have always distinguished expertise from authority.
The smartest lawyer does not automatically become a judge.
The most knowledgeable pharmacologist does not personally approve a drug.
The best accountant does not acquire the authority to certify any company's financial statements simply because of superior analytical ability.
Expertise does not confer authority.
Institutions confer authority.
Why should intelligence automatically confer authority upon machines?
AI can win the intelligence race and still lose the right to decide.
The Intelligence–Trust Gap
This distinction becomes important because intelligence and trust are developing at very different speeds.
AI capability may improve dramatically within months.
Institutional trust rarely does.
New models can be trained, deployed and replaced rapidly. Institutional systems develop through standards, evidence, testing, regulation, accountability, precedent and accumulated experience.
One curve is moving exponentially.
The other is moving institutionally.
The gap between them is the Intelligence–Trust Gap.

The greater that gap becomes, the more capable AI becomes of exercising
consequential power before society has developed adequate mechanisms to delegate that power responsibly.
This does not necessarily mean that people will stop using AI.
In fact, the opposite may happen.
People may increasingly depend upon AI for communication, search, coding, research, healthcare information, financial analysis, education and business decisions while simultaneously becoming more uncertain about whether the systems themselves can be trusted.
Usage is not the same as trust.
Dependency is not the same as legitimacy.
That distinction matters enormously.
Society Is Accumulating AI Trust Debt
I would describe another phenomenon emerging alongside the Intelligence–Trust Gap as AI trust debt.
Trust debt accumulates when society is asked to accept additional capability and autonomy today while adequate accountability, containment, independent validation and institutional oversight are expected to arrive later.
The recent OpenAI and Hugging Face disclosures are important precisely because both organizations publicly examined what happened. OpenAI restricted the affected systems and involved external security organizations.
Hugging Face published a detailed reconstruction of the intrusion.
Transparency matters.
But transparency following an incident is not the same thing as institutional trust before the incident.
Society will still ask:
Who decided what capabilities the AI could access?
Who independently tested the containment environment?
What authority was given to the agent?
Who monitored its actions?
Which boundaries existed only as instructions, and which were actually enforced?
Who decides whether the remediation is sufficient?
And perhaps most importantly:
Can the organization developing an increasingly powerful intelligence also be the final authority deciding whether that intelligence is sufficiently safe?
These are no longer merely technical questions.
They are questions of legitimacy.
This creates a difficult societal problem. People may become uncertain whether technology companies can control increasingly capable AI systems while simultaneously remaining uncertain whether governments and regulators understand the technology well enough to oversee those companies.
That creates a double trust deficit.
Society does not fully trust the creators.
Society does not necessarily trust the institutions governing the creators either.
Under those conditions, every technical incident becomes larger than the incident itself. It becomes evidence in a broader debate about whether anyone is truly in control.
The Collision Between Science and Institutions
This brings us to a deeper collision.

Science and institutions move differently.
Science broadly follows:
Discover → Improve → Replace → Iterate
Institutions broadly follow:
Establish rules → Validate → Standardize → Document → Challenge → Change carefully
Science celebrates discovering that yesterday's model was wrong.
Institutions must ask what happens to all the decisions made using yesterday's model.
AI developers want the model released in June to be materially better than the model released in January.
An institution may need the system operating in June to remain predictably consistent with the system validated in January.
Neither side is irrational.
Science is optimized for discovery and improvement.
Institutions are optimized for continuity, accountability and legitimacy.
AI brings these two operating philosophies into direct collision.
Imagine that an FDA workflow is validated using one model.
Three months later, a substantially better model becomes available.
Should the institution continue using the inferior but validated model?
Should it immediately upgrade?
Does the entire workflow require revalidation?
Should regulators validate the model or the process surrounding the model?
What happens when a hosted model changes without the institution fully understanding what changed internally?
What happens when the new model is statistically more accurate but behaves differently on edge cases?
The AI researcher naturally asks:
What can the new model do that the previous model could not?
The institution must ask an equally important question:
What else changed?
The second question may ultimately govern institutional adoption more than the first.
Prompts Express Intention. Architecture Enforces Authority
The current AI development philosophy often assumes that governance can be placed around a capable model after the intelligence has been created.
First build the most powerful model.
Then add guardrails.
For many low-risk applications, this may be sufficient.
For institutional AI, it may be backwards.
The key distinction is between asking intelligence to respect a boundary and building a system in which the boundary cannot be crossed without authorization.
Suppose an AI is instructed:
Do not transfer more than $10,000 without approval.
That is an instruction.
Now suppose the execution system technically prevents any transaction above $10,000 unless a separately authenticated human or system grants authorization.
That is a control.
Suppose an AI agent is told not to examine certain confidential information.
That is an instruction.
If the agent's identity and permissions make that information inaccessible, that is architecture.
Suppose the model is asked to maintain a record of everything it does.
That requires trusting the model.
If the surrounding infrastructure independently records every API request, tool call, data access, authorization and action, the institution no longer needs to rely solely upon the model's own account of its behavior.
Prompts express intention. Architecture enforces authority.
Human institutions learned this distinction centuries ago.
We do not simply tell people to behave correctly.
We separate duties.
We restrict access.
We require signatures.
We establish approval limits.
We preserve original records.
We independently reconcile transactions.
We create escalation procedures.
We allow decisions to be appealed.
AI will require the digital equivalent.
This is why institutional trustworthiness cannot reside entirely inside the model.
It must be engineered around it.
A Chain of Custody for Machine Intelligence

In 2020, I wrote that intelligence is not factual like data.
Two intelligent systems can examine the same information and reach different conclusions. A model can be highly confident and still be wrong. As enterprises increasingly relied on machine intelligence, I argued that they would eventually need a "single source of truthful intelligence."
At that time, the question appeared primarily organizational.
How should enterprises manage multiple models, competing conclusions, privacy, changing algorithms and model accountability?
The problem has now moved beyond the enterprise.
We are no longer asking only:
Can we trust an AI-generated answer?
We are increasingly asking:
Can we trust an AI-generated action?
That requires something more than model governance.
It requires what might be thought of as a chain of custody for machine intelligence.
Which exact information entered the system?
Where did that information originate?
Was it transformed?
Was OCR or another generative system involved?
Which model and version acted?
What instructions did it receive?
Which tools could it access?
Which actions did it attempt?
Which actions were prevented?
Who authorized consequential actions?
What evidence supported the conclusion?
Can the complete process be reconstructed later?
Can another person or institution challenge the result?
Can the decision be reversed?
This is not conventional logging.
It is the machinery through which machine-generated reasoning and action become institutionally admissible.
AI Must Be Contestable
There is another characteristic of institutions that deserves far more attention in AI design.
Institutional decisions can be challenged.
Court judgments can be appealed.
Financial transactions can be disputed.
Scientific results can be replicated.
Regulatory decisions can be reviewed.
Accounting treatments can be questioned.
Institutions do not earn trust because they are always correct.
They earn trust because they have mechanisms for detecting, examining and correcting mistakes.
This may provide one of the most important requirements for institutional AI:
A trustworthy AI system must not merely produce an answer. It must produce an answer that can survive challenges.
Explainability may help, but explanation alone is insufficient.
The evidence needs provenance.
The action history needs reconstruction.
The model identity needs preservation.
The authority needs to be explicit.
And there must be a process through which another party can contest the conclusion.
Trustworthy AI therefore requires something deeper than explainability.
It requires contestability.
The Emergence of Two-Speed AI
The future may divide into two distinct forms of AI.
Frontier AI will optimize for maximum intelligence, autonomy, exploration, scientific discovery and rapid improvement.
It will remain enormously important. Progress in science, software development, cybersecurity, engineering and research may increasingly depend upon it.
Institutional AI, however, will optimize for something different: bounded authority, stable contexts of use, controlled access, preserved evidence, independent validation, human accountability and contestability.
Institutional AI may sometimes use an older model.
It may be narrower.
It may be slower.
It may be prevented from doing things that the underlying model is technically capable of doing.
From the perspective of pure intelligence, these constraints may appear inefficient.
From the perspective of institutional trust, the constraints are the product.
Commercial aviation does not certify an aircraft merely because it is the fastest aircraft ever built.
Courts do not admit evidence simply because it was produced by the smartest available expert.
The FDA does not approve a therapy because its scientific mechanism appears intellectually compelling.
Capability matters.
But admissibility requires evidence.
AI will face the same distinction.
Will Intelligence Lose to Trustworthiness?

Intelligence will not lose the scientific race.
The incentives for creating increasingly capable AI are simply too strong.
Companies, governments and researchers will continue pushing the frontier.
Intelligence may not lose in consumer markets either. Humans routinely use technologies that they do not completely trust when their utility is sufficiently high.
But the more important battle will occur wherever AI seeks consequential authority.
Healthcare. Banking. Insurance. Courts. Regulation. Critical infrastructure. Corporate controls. Defense. Government.
In these environments, trustworthiness can become a veto condition.
A somewhat less capable system that is bounded, reproducible, auditable and accountable may be institutionally superior to a more intelligent system whose behavior cannot reliably be constrained or reconstructed.
That leads to a counterintuitive possibility:
The most intelligent AI may not become the most institutionally important AI.
The winning institutional system may instead be the most capable intelligence that can be governed—not the most capable intelligence that exists.
The Real AI Race
For much of the last decade, the defining question has been:
Who will build the most intelligent model?
The next decade will introduce another question:
Who will build intelligence that society is willing to authorize?
The first race is about capability.
The second is about trust and legitimacy.
There is no reason to believe that these races will proceed at the same speed.
AI capability advances through computation, algorithms, data, capital and intense competition.
Institutional trust advances through evidence, standards, independent verification, accountability, legal structures, precedent and accumulated experience.
One can advance in months.
The other often requires years.
The central risk of the next phase of AI may therefore not be that intelligence stops progressing.
It may be precisely the opposite.
Intelligence may progress so rapidly that trustworthiness and institutional legitimacy cannot keep pace.
If that happens, we may enter a strange world in which society possesses extraordinary machine intelligence, uses it everywhere, increasingly depends upon it—and still refuses to trust it with its most consequential decisions.
Can AI win institutional trust?
Yes.
But not merely by becoming more intelligent.
AI will ultimately have to accept the same conditions that human civilization imposed upon human intelligence: boundaries, evidence, oversight, accountability and the right to challenge decisions.
An AI answer must not merely appear correct.
It must survive scrutiny.
An AI action must not merely accomplish an objective.
It must remain within legitimate authority.
And an AI system must not merely tell us that it can be trusted.
It must create the evidence through which that trust can be independently established.
Intelligence earns attention. Trustworthiness earns deployment. Legitimacy earns authority.
The real AI race may therefore not be intelligence versus intelligence.
It may be intelligence versus the speed at which trust can be built.
And if the Intelligence–Trust Gap grows too wide, AI may not lose the race for intelligence.
It may lose something considerably more important:
society's permission to act.
About The Author

Vivek Gupta is Founder & CEO at SoftSensor.ai, specializing in Enterprise AI Development and Adoption. For more insights on AI’s impact on process design and organizational behavior, feel free to reach out at vivek.gupta@softsensor.ai.





Comments