TL;DR
Companies buy AI the way they buy anything else: pick the best supplier, keep a second one for safety. With models that logic quietly fails. Research from Cornell shows that models from different vendors and different architectures tend to be wrong about the same things, and that the more accurate two models are, the more their errors line up. A large study of four million job applications shows what that does to the people on the receiving end. The risk nobody puts a number on is not accuracy. It is agreement.
Twenty firms. Twenty different models, each drawn at random from a different vendor. Around one in five candidates was still rejected by all twenty.
That result sits in a Cornell paper from June 2025 and I have yet to hear it quoted in a single enterprise AI meeting.
The assumption nobody writes down
Every risk model in business runs on independence. Two reviewers catch more than one. Two suppliers survive what one supplier does not. Two opinions are better than one opinion heard twice. We never state it, which is exactly why we never check it.
In June 2026 a team from Stanford and Northeastern published the largest audit of hiring algorithms I know of: 4,197,168 applications from 3,372,132 applicants, 1,746 positions, 156 employers, all screened by one vendor's algorithms between December 2018 and December 2022. About 4 percent of applicants who applied to ten positions were rejected by all ten. Under independence that number should collapse as you apply more widely. It does not. To reach a near certain chance of one single human reading your file, you would need to apply to 25 positions, not 10.
The obvious fix is to spread the work across vendors. The Cornell group tested that assumption across 420 model evaluations from more than a dozen providers. When two models are both wrong, they agree on the wrong answer about 60 percent of the time on one benchmark, against a chance baseline of 33 percent. On another, 42.3 percent against 12.7. Practically every pair in the sample sits above its baseline.
Different logo. Same blind spot.
The sentence that should worry a procurement team
Buried in the regression tables is the part that changes how you buy: the more accurate two models are, the more their errors correlate. Pick the best model on the leaderboard, as everyone does, and you move toward the part of the distribution where mistakes overlap most. Every buyer optimises their own decision correctly, and the system gets more fragile with each rational choice. There is no column for that on a vendor scorecard.
The same shape, from the buyer's chair
The industry has started to notice the symptom without naming the cause. A Cloud Security Alliance paper from 19 June 2026 reports that 91 percent of executives lack full visibility of their AI dependencies, 71 percent say switching their main model vendor would be hard, and 81 percent expect severe or critical disruption from a seven day outage at one provider. High signal disruption days went from 6 in the first quarter of 2025 to 51 in the first quarter of 2026.
On 6 October 2026 an IT executive at Citi gave it a decent name in Forbes: cognitive concentration. Many applications, one underlying mind, correlated judgments. Her framing treats it as a dependency problem, solvable by diversifying. The Cornell number says the harder half out loud: diversifying the vendor does not buy you independence back.
Nineteen seventy
American farmers had dozens of seed brands to choose from. Roughly 90 percent of the hybrid corn planted carried the same Texas male sterile cytoplasm, because it made hybrid seed cheaper to produce. Southern corn leaf blight found that one shared trait and took around 15 percent of the corn belt's crop in a single season.
The fields looked diverse. The seed was not.
Why this matters for your business
Two questions belong in your next AI review, and neither is about accuracy. First: where do we assume two systems are checking each other when they share a training lineage? Credit decisions, fraud flags, CV screening, supplier risk, content moderation. Second: is our fallback a different model, or a different invoice?
And a warning about the comfort of human oversight. A reviewer who sees three systems agree does not experience three independent judgments. They experience consensus, which is the most persuasive form a correlated error can take.
Your second vendor is certainly a second contract. Whether it is a second opinion is an empirical question, and almost nobody is measuring it.