1.2 seconds: anatomy of an accountability failure
A claims system produced hundreds of thousands of denials at an average of 1.2 seconds each — with a doctor's name on every one. The lesson is not that automation is bad. It is that a system can produce the evidence of human review without producing human review.

In March 2023, ProPublica published a story about an internal Cigna claims review system called PxDx. The number that stuck was 1.2 seconds.
According to internal company records reviewed by ProPublica and The Capitol Forum, Cigna doctors denied more than 300,000 requests for payment over a two month period in 2022 using the system, spending an average of 1.2 seconds on each case. One medical director was recorded as denying roughly 60,000 claims in a single month.
Think about 1.2 seconds. That is barely enough time to see that something appeared on a screen. It is not enough time to open a medical record, understand the patient's circumstances, compare those circumstances against coverage criteria, exercise medical judgment, and make an individual decision.
And yet, from the outside, many of the artifacts of a legitimate process were there.
There was a claim.
There was a rule.
There was software.
There was a physician.
There was a decision.
There was a timestamp.
There was a doctor's name attached to the outcome.
There was a record showing that the process had happened.
That is exactly why this case matters.
The interesting part is not simply that software was used. Software should be used. AI should be used.
Insurance companies process enormous volumes of information. Banks do too. Payment companies do too. Hospitals do too. Governments do too. Nobody seriously wants every routine decision manually processed by a person with a clipboard.
The interesting part is much narrower. The system was capable of producing evidence that a human had reviewed a decision without necessarily producing meaningful human review.
That distinction is going to become one of the defining problems of the AI economy. Because as AI becomes better at doing the work, organizations are going to become increasingly tempted to preserve the human only as a signature.
And a signature is not the same thing as judgment.
The problem was not automation
There is an easy version of this story that says automation is bad and humans are good. That is not the lesson. In fact, that lesson would probably make the system worse.
Imagine forcing a physician to manually inspect every routine insurance claim that software could safely process. It would be expensive, slow, frustrating for doctors, frustrating for patients, and probably less consistent.
Good automation removes repetitive work.
Good AI compresses information.
Good models identify obvious cases.
Good software routes exceptions.
The problem begins when an organization says there is human judgment at a particular point in the workflow when, operationally, there is not.
ProPublica reported that Cigna's system compared diagnosis codes against procedures and tests, then presented cases to medical directors who could process large batches of denials. Former Cigna employees told the publication that doctors could sign off without opening individual patient files. Cigna disputed ProPublica's characterization of the process, describing the reporting as biased and incomplete and saying PxDx was designed to accelerate processing of certain routine claims.
That disagreement is important. We do not need to resolve every disputed fact about Cigna to learn something from the structure. The structural question is simpler.
What does it actually mean to say a human reviewed an AI assisted decision?
That question matters far beyond Cigna.
Human in the loop can mean almost anything
Human in the loop has become one of those phrases that sounds precise until you ask what it means.
A human clicked approve.
Human in the loop.
A human was copied on an email.
Human in the loop.
A supervisor can theoretically intervene.
Human in the loop.
A doctor signed 500 decisions generated by software.
Human in the loop.
An employee had five seconds to disagree with an algorithm.
Human in the loop.
A licensed professional reviewed the underlying evidence, understood the case, disagreed with the AI where appropriate, wrote a reason, signed the decision, and accepted responsibility for it.
Also human in the loop.
Those are obviously not the same thing. Yet many governance documents treat them as if they are.
That is the problem.
Presence is not oversight. A person being somewhere inside the process does not automatically create accountability. Meaningful human oversight requires more.
The arithmetic often tells you the truth before the audit does
One useful lesson from the Cigna story is that sometimes you do not need a sophisticated forensic investigation. You can start with arithmetic.
300,000 decisions.
Two months.
A limited number of medical directors.
1.2 seconds per decision on average.
At some point the throughput itself tells you what kind of activity could realistically have occurred.
This is useful because organizations tend to evaluate controls qualitatively.
Do we have human review?
Yes.
Is the reviewer licensed?
Yes.
Is the decision logged?
Yes.
Do we retain the records?
Yes.
Every box gets checked.
But one quantitative question can expose the entire thing.
Could the claimed review actually have happened at this volume?
That is a different question. And it is a much better one.
If a company says a person reviews 2,000 complex decisions every hour, the organization should not need a whistleblower to discover there is a problem. The math is already the audit finding.
A control can exist and still not operate
Auditors have dealt with versions of this problem forever.
A policy can exist without being followed.
An approval field can exist without anybody investigating what they approve.
A manager can sign a document without reading it.
A committee can technically review something that nobody had enough information to evaluate.
There is a difference between the existence of a control and the operating effectiveness of a control. AI makes that old problem much more important because AI radically increases throughput.
Before automation, fake review had a natural speed limit. A human could only rubber stamp so many things.
AI changes the economics.
The machine can prepare ten thousand decisions.
The human can technically approve ten thousand decisions.
The database can record ten thousand human approvals.
Now the organization has ten thousand perfect looking audit records. The problem is that the evidence may describe an approval event, not a judgment event.
That distinction sounds philosophical. It is not. It determines whether the human actually changed the risk of the system.
The click is real, the review may not be

This is why ordinary system logs are weaker evidence than they appear. Suppose a log contains this:
Reviewer: Dr. Smith
Decision: Denied
Time: 10:42:31
User authenticated: Yes
That record can be completely accurate.
Dr. Smith really logged in.
Dr. Smith really clicked deny.
The timestamp really is 10:42:31.
Nothing was fabricated.
But the record still does not answer the important questions.
What information did Dr. Smith actually see?
Did Dr. Smith inspect the underlying evidence?
How much time was available?
Was the recommendation presented neutrally, or was the interface designed around confirmation?
Was there a meaningful option to disagree?
Would disagreement create extra work?
Could the reviewer request additional information?
Was the reviewer qualified for this particular issue?
Was the reviewer financially or operationally pressured toward one outcome?
Did the reviewer know that thousands of additional cases were waiting?
Could the reviewer reverse the algorithm without needing somebody else's permission?
The log normally does not know.
A click log proves a click. It does not prove judgment.
The most important question is whether the human could realistically say no
This is the question we think people should ask much more often.
Could the human realistically disagree with the machine? Not theoretically. Realistically.
An interface can contain a reject button while making rejection practically impossible.
Maybe accepting the recommendation takes one click while disagreeing requires twelve fields.
Maybe rejecting the AI creates a supervisor escalation.
Maybe the reviewer is evaluated on throughput.
Maybe a queue grows every time they investigate something.
Maybe the AI recommendation is displayed first in large text while contradictory evidence is hidden three screens away.
Maybe the reviewer receives no underlying documents at all.
Maybe the decision is already effectively made before it reaches them.
In each case, there is technically a human. But the human has become part of the automation.
That is not necessarily oversight. It may simply be another processing step.
A genuine checkpoint needs the possibility of changing the outcome.
If disagreement never occurs, that does not automatically mean the AI is excellent. It might mean the human is ornamental.
One of the first metrics we would want to see in a serious review system is disagreement.
How often did humans overturn the model?
How often did they ask for more information?
How often did they escalate?
How often did they modify the recommendation?
What happened after disagreement?
A reviewer who never changes anything may be reviewing a perfect system. Or may not be reviewing at all.
Those possibilities should not look identical in the data.
Qualification matters too
There is another mistake hidden inside the phrase human oversight. It treats humans as interchangeable. They are not.
If the decision is whether a medical service is clinically necessary, putting a random employee into the workflow does not solve the problem.
If the question involves a legal opinion, credentials matter.
If an insurance decision requires a licensed adjuster, credentials matter.
If an accounting assumption needs professional review, expertise matters.
If an AML alert needs escalation, domain knowledge matters.
The human must not merely exist. The human has to be capable of owning the specific decision.
California has made that concept unusually explicit.
Under SB 1120, an AI system may assist utilization review, but it may not deny, delay, or modify health care services based on medical necessity. A medical necessity determination must be made by a licensed physician or licensed health care professional competent to evaluate the specific clinical issues, and that person must consider the provider recommendation, the individual's medical history, and the individual's clinical circumstances. The law also makes AI systems used in this process available for regulatory audit or compliance review.
That is much more specific than saying put a human in the loop. It asks which human.
Individualization matters
The second principle is individualization.
AI is incredibly good at patterns. That is why we use it. It can inspect millions of prior cases and identify a pattern no person could see.
But a population level pattern and an individual decision are not always the same thing.
Imagine a model saying that 97 percent of cases with a particular combination of codes do not qualify. That can be extraordinarily useful. It can tell the organization where to focus attention.
But the person in front of you might be in the other 3 percent.
In high consequence decisions, the important question is often not whether the model identified the common case. The important question is whether this is the uncommon case.
California's statute explicitly requires consideration of individual clinical circumstances and says the tool cannot base its determination solely on a group dataset.
Maryland has now adopted a closely related approach. Its insurance law requires AI used in utilization review to consider individual medical or clinical information, prohibits the tool from replacing the health care provider's role, makes the system open to regulatory inspection, and states that an AI tool may not deny, delay, or modify health care services.
Again, the direction is clear.
Use the machine for what machines are good at. Do not pretend a statistical pattern is the same thing as individualized professional judgment.
Texas goes even further
Texas enacted SB 815 in 2025.
The statute says a utilization review agent may not use an automated decision system to make, wholly or partly, an adverse determination. It also gives the insurance commissioner authority to audit and inspect the use of automated decision systems in utilization review.
The law took effect September 1, 2025, with the relevant requirements applying to utilization review for health benefit plans delivered, issued, or renewed beginning January 1, 2026.
Notice what keeps appearing in these rules.
A responsible person.
Individual circumstances.
A reason.
Regulatory inspection.
Evidence.
Those concepts are starting to form an accountability architecture.
Insurance is not waiting for one universal AI law
The regulatory direction is broader than a few state statutes.
The National Association of Insurance Commissioners adopted its Model Bulletin on the Use of Artificial Intelligence by Insurance Companies in December 2023.
The NAIC says insurers remain responsible for legal and regulatory compliance when AI supports their decisions, and regulators may ask for information about how those systems are governed and used. The NAIC's current work includes an AI Systems Evaluation Tool intended to help regulators examine AI use, governance, risk mitigation, higher risk models, and data inputs. As of March 2026, the tool was being piloted by 12 states.
This matters because companies sometimes imagine there will be one future AI law telling them exactly what database field to store. That is unlikely to be how accountability develops.
Instead, existing obligations around fairness, professional responsibility, consumer protection, explanations, records, internal controls, and regulatory examination are being applied to increasingly automated systems.
The technology is new. The question regulators ask is ancient.
Show me what happened.
Banking has the same problem
Insurance is simply a vivid example. Credit decisions create the same structural issue.
A lender can use an extraordinarily complex model. That does not eliminate the obligation to explain an adverse decision.
The Consumer Financial Protection Bureau has said creditors using AI and other complex models still have to provide accurate and specific reasons for adverse actions. Companies cannot hide behind the complexity of their models or use a generic checklist when it does not reflect the real reason a consumer was denied.
That creates an interesting chain of accountability.
The model produces a recommendation.
The organization takes an action.
A person may review it.
The customer receives an explanation.
A regulator may later ask whether that explanation actually corresponds to the decision.
Every step creates another place where accountability can become fictional.
A model explanation can be fictional.
A human review can be fictional.
A reason code can be fictional.
A log can be accurate while the story it seems to tell is false.
This is why AI governance cannot only be about model accuracy. It has to be about decision integrity.
Fraud systems have the same problem
Consider a payment platform.
A model flags a merchant account as suspicious. Maybe the model is 99 percent accurate. Great.
But the remaining one percent might contain businesses whose livelihoods depend on access to those funds. So the company puts a human reviewer in the loop.
Now imagine the reviewer gets 400 alerts an hour.
Each one arrives with a giant red label saying HIGH RISK.
The default action is freeze.
The evidence supporting the algorithm appears immediately.
Evidence contradicting it requires opening three additional systems.
The reviewer is evaluated on cases closed per hour.
Technically there is human review. Operationally there is confirmation bias implemented as software.
The existence of the human did not necessarily reduce the model's risk. The workflow may actually amplify it.
Hiring has the same problem
An AI ranking system scores candidates. A recruiter receives the top candidates. The recruiter technically makes the final choice.
That sounds like human oversight.
But what happened to candidates ranked number 900?
Did anybody see them?
Could the recruiter realistically question the ranking?
Was the human choosing candidates, or choosing from a world already constructed by the model?
This is another reason the final click is a poor place to locate accountability. A decision system can constrain human judgment long before the final decision screen appears.
Legal work has the same problem
An AI drafts a contract. A lawyer approves it. The document goes out.
The organization records that a licensed lawyer reviewed the contract. Fine.
But now ask what reviewed means.
Did the lawyer compare the draft against the client's actual requirements?
Did they inspect the changed clauses?
Did they know which text was generated?
Were relevant jurisdictions identified?
Was the lawyer responsible for the matter?
Could they edit or reject the document?
Was the reviewed version the same version that was later executed?
That last question is particularly important.
You can have completely legitimate human review followed by an accountability failure if the reviewed artifact and the final artifact are different. The evidence therefore has to bind the person not just to a general workflow, but to the thing they actually reviewed.
The coming problem with AI agents
All of these examples become more important with agents.
Traditional software waits for instructions. Agents increasingly perform sequences.
Read the email.
Open the claim.
Check the policy.
Query the database.
Generate the recommendation.
Create the payment.
Send the notice.
Update the CRM.
Close the case.
The more of that sequence the AI performs autonomously, the more important it becomes to define exactly where a human decision is required.
Not everywhere. That would defeat much of the point of automation. At the right place.
A good system might autonomously execute 95 percent of a workflow and pause for a person only when a decision crosses a risk threshold. But when that pause occurs, it has to be real.
Otherwise we have built a very sophisticated autonomous system with a ceremonial human attached to it.
So what does meaningful human review actually require?
We think there are at least six things.
First, identity.
Who made the decision?
Not which department.
Not which generic service account.
Not which shift.
Which person.
Second, qualification.
Why was this person competent or authorized to decide?
Their role may be enough for some decisions. For others, you need a professional credential, license, jurisdiction, specialty, or other verified qualification.
Third, the actual evidence.
What did the reviewer receive?
What version did they see?
What was the AI recommending?
What supporting information was available?
What policy or criteria applied?
A decision without the underlying context is difficult to evaluate later.
Fourth, authority.
Could this person genuinely change the outcome?
Could they approve?
Deny?
Overturn?
Request more information?
Escalate?
Stop the workflow?
A human who lacks decision authority is not necessarily a decision maker.
Fifth, an individualized decision.
What did the reviewer actually decide, and why?
Not simply APPROVE. A structured reason matters.
The reason does not have to be a thousand word essay. It does have to tell us something about the judgment exercised.
Sixth, evidence integrity.
Can the record later be altered?
Can one decision record be substituted for another?
Can the organization change the reviewer field after the event?
Can somebody regenerate a nicer record after receiving a regulatory request?
If all accountability evidence sits inside a database controlled by the organization being audited, the auditor still has to trust that organization.
That may be perfectly acceptable in many low risk situations. In higher consequence workflows, we can do better.
What a real accountability record should look like
Imagine an AI recommends denying a claim.
Before the denial leaves the system, the workflow pauses.
A qualified reviewer receives the exact recommendation and relevant supporting material. The system verifies that the reviewer has the credential required for that type of decision.
The reviewer can approve the recommendation, overturn it, request additional information, or escalate it. They make a structured decision.
The system records their identity, qualification, decision, relevant evidence reference, and time. The decision becomes bound to the specific workflow instance. A tamper evident fingerprint is created. The record can later be independently checked.
Now we have something much more interesting than:
User 873 clicked approve.
We have a decision artifact.
That artifact does not magically prove the human thought deeply. No cryptographic hash can prove cognition.
That distinction is important. Technology should not claim to prove something it cannot prove.
What it can do is remove ambiguity around identity, authorization, artifact integrity, decision lineage, and record alteration. Then operational controls can address workload, review quality, minimum evidence, sampling, disagreement rates, and performance.
That combination is much stronger than a log.
This is the distinction between logging and proving
Logging asks:
What did our system record?
Proof asks:
What can somebody else verify?
Those are different questions.
Most organizations today operate almost entirely in the first category.
Their application creates the decision.
Their application creates the log.
Their database stores the log.
Their administrators control the database.
Their reporting system produces the evidence.
Then, when somebody asks what happened, the organization exports a PDF from its own system.
That is normal. It is also increasingly insufficient for decisions where accountability itself matters.
A stronger model separates the decision from the evidence that the decision occurred.

This is the problem HumanAgent is built around
HumanAgent is based on a simple idea.
AI should automate aggressively. Humans should not be dragged into every routine task.
But when a workflow reaches a point where judgment, accountability, trust, or professional responsibility is actually required, the human checkpoint should be explicit.
HumanAgent lets an AI workflow call a human checkpoint through an API. The task can specify the type of reviewer required, the service level, the budget, and the context needed to make the decision. The result comes back in structured, machine readable form so the automated workflow can continue.
The reviewer can also come from the company's own internal team. HumanAgent's current model supports internal first routing, where the organization supplies the person and HumanAgent supplies the checkpoint and record, or the review can be routed to a verified external operator.
The point is not to create a marketplace of people clicking buttons. That would recreate the problem.
The point is to make the human decision a first class part of the system architecture.
Identity matters.
Credential matters.
Decision matters.
Record integrity matters.
Why we call it Proof of Human
Every completed HumanAgent checkpoint produces a record tied to a reviewer and credential.
HumanAgent creates a SHA 256 fingerprint from canonical decision fields, and completed decision records can be sealed externally. The current sealing model takes a fingerprint of decisions completed during a period and publishes that fingerprint somewhere outside HumanAgent's control. A later alteration would therefore cause the recomputed fingerprint to stop matching the published seal.
The important distinction is that this is not supposed to prove a philosophical concept like attention. It proves a narrower set of things much better.
The record existed.
The record is attributed.
The decision is associated with a particular reviewer and credential.
The stored record matches the fingerprint.
The record has not silently changed since it was externally sealed.
That is much closer to the kind of evidence an auditor, insurer, regulator, customer, or counterparty can independently inspect.
And it moves accountability away from:
Trust our database.
Toward:
Verify the record.
But cryptography alone is not accountability
This is worth repeating because the technology industry loves technical solutions to human problems.
A perfect hash does not fix a terrible review process.
You can cryptographically prove that somebody rubber stamped something. Congratulations, now you have tamper evident rubber stamping.
The record layer and the workflow layer have to work together.
If the person does not have enough information, fix the workflow.
If the person does not have authority, fix the workflow.
If the reviewer has 8,000 cases waiting, fix the capacity model.
If the interface manipulates the reviewer toward accepting the AI recommendation, redesign it.
If credentials are irrelevant to the decision, do not pretend they solve it.
If disagreements are punished, the review is compromised.
If the reviewer cannot see the thing they are supposedly reviewing, stop calling it review.
Proof does not replace governance.
Proof makes governance inspectable.
The best evidence of meaningful oversight may be disagreement
One metric deserves much more attention. Human disagreement with AI.
Organizations often celebrate agreement rates.
Our AI recommendation matched the human 99.8 percent of the time.
Maybe that is excellent. But before celebrating, ask another question.
Was disagreement easy?
A healthy human checkpoint should occasionally produce outcomes such as:
Overturn.
Request more information.
Escalate.
Modify.
Reject recommendation.
Unable to determine.
Those outcomes are evidence that the human is doing something the machine cannot completely predetermine.
In HumanAgent's own healthcare workflow, for example, the intended role of a clinician reviewing a medical necessity recommendation includes the ability to confirm or overturn the recommendation after reviewing the clinical record and relevant criteria.
That is a much better design target than simply proving a person appeared somewhere in the process.
Accountability has a cost
There is another uncomfortable part of this conversation.
Real review costs money. It takes time.
Qualified people cost more than generic labor.
An experienced physician costs more than a button.
A licensed attorney costs more than an automated approval.
A qualified adjuster costs more than a generic contractor.
This creates economic pressure to make the human step as thin as possible.
That pressure is not evil. It is normal. Companies optimize.
The question is whether the optimization crosses the point where the control no longer performs the function used to justify its existence.
The goal should therefore not be maximum human involvement. That would be inefficient.
The goal should be minimum sufficient human involvement.
Automate everything that can responsibly be automated. Then make the remaining human decision real.
The future is probably not human versus AI
It is AI, human, evidence.
AI does what it does best.
It reads enormous quantities of information.
It finds patterns.
It drafts.
It predicts.
It scores.
It compares.
It retrieves.
It summarizes.
It handles the common case.
The human does what we still need humans to do.
Exercise judgment.
Take responsibility.
Recognize an exception.
Interpret ambiguous context.
Apply professional standards.
Say no.
And then the evidence layer answers the question that appears months or years later.
Who decided?
Under what authority?
Using what information?
What did they decide?
Can you show me?
That last layer is going to matter more than many companies expect.
This is not only about compliance
There is a temptation to put all of this under compliance. That would be a mistake.
There are plenty of decisions where no regulation forces a human checkpoint, but a business may still want one.
A company may want human approval before sending a seven figure payment.
A founder may want a person to review an acquisition model before using it in negotiations.
An AI coding agent may be able to deploy software automatically, but a company may still require a qualified engineer to approve a security sensitive change.
A procurement agent might negotiate automatically but require a person to approve contracts above a threshold.
An AI sales system may generate thousands of messages but require a human to review communications involving regulated claims.
The architecture is the same.
Automation.
Threshold.
Human judgment.
Evidence.
Continue.
This is not an anti AI architecture. It is what allows more automation to be trusted.
Better human checkpoints can actually increase automation
That may sound counterintuitive.
If you make human oversight stronger, do you not slow the AI down?
Sometimes.
But the alternative is often that companies refuse to automate the workflow at all.
Suppose a company wants to automate 90 percent of a regulated process but cannot reliably control the remaining 10 percent. The risk team may block the entire project.
A robust checkpoint changes that conversation.
The company can say:
Routine cases proceed automatically.
Higher risk cases pause.
Specific credentials are required.
The reviewer owns the decision.
The record is retained.
The decision can be verified.
Now the organization may be comfortable automating far more of the workflow.
Human accountability is not necessarily a brake on AI. It can be infrastructure for deploying more AI.
The Cigna lesson is therefore bigger than Cigna
The most interesting thing about the 1.2 second number is not that it makes for a shocking headline. It gives us a very clean example of a general principle.
Accountability can be simulated by artifacts.
A person's name can appear on a decision without telling us how meaningful their involvement was.
A log can be accurate and still tell an incomplete story.
A policy can be followed formally while failing operationally.
A human can be in the loop while having almost no practical influence over the outcome.
And a company can accumulate millions of records proving activity without proving judgment.
That is the accountability problem AI makes urgent.
Three questions every high stakes AI workflow should answer
If you operate an AI system that affects money, health, access, legal rights, safety, employment, insurance, or another person's significant interests, start here.
Who decided?
Were they qualified and genuinely able to decide otherwise?
Can you prove what happened?
If any answer is fuzzy, the system probably has an accountability gap.
Not necessarily an AI problem. An accountability problem.
And those are harder to fix after the fact.
What is meaningful human review?
Meaningful human review is not simply the presence of a person somewhere inside an automated workflow. It means a person with appropriate authority and, where necessary, qualifications receives enough relevant information to independently evaluate the decision, has a realistic ability to change the outcome, and creates a record attributable to that judgment.
What is human in the loop AI?
Human in the loop AI is an architecture where a human participates in some part of an automated process. The term by itself says almost nothing about the quality of that participation.
The useful questions are where the human enters, what information they receive, what authority they have, what qualifications they need, and what evidence survives afterwards.
Does putting a human in the loop make an AI system compliant?
No.
Human involvement may be one requirement among many. Regulated systems can also face obligations concerning data, fairness, discrimination, explanations, privacy, record keeping, professional qualifications, individual circumstances, and regulatory access. A ceremonial human does not cure a defective system.
Why are ordinary audit logs not enough?
Sometimes they are enough.
But a normal application log usually proves what the application recorded. It may not establish that meaningful review occurred, that the reviewer was appropriately qualified, that the reviewed artifact was the final artifact, or that the record was not later changed.
For higher consequence decisions, the evidence architecture should be designed around the question an independent third party will eventually ask.
How do I know?
What is Proof of Human?
In HumanAgent, Proof of Human is the verifiable record created around a completed human checkpoint. It associates the decision with a reviewer and credential, creates a cryptographic fingerprint of canonical record fields, and supports external sealing so the integrity of the record can later be independently checked.
It does not claim to read a person's mind. It makes the accountable parts of the interaction harder to fake, substitute, alter, or leave ambiguous.
Where should companies use human checkpoints?
Not everywhere.
Use them where the cost of an unreviewed automated decision becomes larger than the cost of review. That can include adverse insurance decisions, medical necessity review, underwriting exceptions, large financial transfers, suspicious transaction escalations, contract approvals, regulated communications, high risk model outputs, safety decisions, and other workflows where a named person still needs to own the judgment.
HumanAgent currently supports checkpoint patterns across insurance, healthcare, banking and payments, legal and compliance, finance, and other operational workflows.
The question companies should ask now
The important question is not:
Do we use AI?
Almost everyone will.
It is not:
Do we have humans?
Almost everyone does.
And eventually it will not even be:
Do we have logs?
Of course you have logs.
The harder question is:
If this decision is challenged two years from now, what exactly will we be able to prove?
Can you identify the person?
Can you establish their qualification at the time?
Can you identify the exact decision they reviewed?
Can you show what the AI recommended?
Can you show what the person decided?
Can you establish that they had authority to disagree?
Can you show that the record you are presenting today is the same record that existed then?
Can somebody outside your organization verify any of it?
That is the standard companies should start designing toward.
Not because every regulator has already written it into one neat rule. They have not.
Because it is the natural consequence of putting increasingly autonomous systems into increasingly consequential workflows.
The more decisions machines make, the more valuable genuine human judgment becomes at the few places where it still matters. And the fewer human decisions we make, the more important it becomes to prove that those decisions were real.
1.2 seconds gave us a useful warning. Do not confuse the evidence that somebody clicked with evidence that somebody decided.
The next generation of AI infrastructure needs to know the difference. HumanAgent is being built for that difference.
Sources and further reading
ProPublica and The Capitol Forum, reporting on Cigna's PxDx claims review process, March 2023.
California SB 1120, health care coverage and utilization review requirements for AI and software tools.
Texas SB 815, automated decision systems and adverse determinations in utilization review.
Maryland Insurance Code, utilization review and artificial intelligence requirements.
National Association of Insurance Commissioners, Artificial Intelligence insurance regulatory work and Model Bulletin.
Consumer Financial Protection Bureau, guidance on specific reasons for adverse credit decisions involving AI and complex models.
HumanAgent product documentation and Proof of Human architecture.