
Most AI does not say so
By June 2026, a public database had documented 1,598 court cases worldwide in which AI-fabricated content, including invented citations, fake quotes, and rulings that never existed, was filed with a court, and the tool involved was almost always a consumer chatbot the professional already had [1]. In the founding case, a lawyer confronted about fake citations asked the chatbot to confirm they were real, and it confirmed its own fabrications [2]. The pattern is not confined to law. An audit of 2.5 million biomedical papers, published in The Lancet in May 2026, found fabricated references in roughly 2,800 of them, at a rate that has risen twelvefold since 2023 [3]. In a global survey of clinicians across 15 specialties, 91.8 percent reported encountering AI hallucinations, and 84.7 percent believed those errors could cause patient harm [4]. Companies including Pfizer, Novartis, Sanofi, GSK, and AstraZeneca, five of the largest in the industry, have moved generative AI from pilots to enterprise scale in medical affairs, and their medical affairs teams must now demonstrate the accuracy and sourcing of AI-assisted work before it is used [6]. Three of those five companies use SteepRock Market Development Services, including Embedded Intelligence, and we see that requirement emerging everywhere. Every chatbot answer arrives fluent and confident, whether it is right or wrong. What follows is what it looks like when an AI checks its own number, fails it, and says so.
The number that did not add up
In August 2026, a global pharmaceutical company preparing for a potential launch asked SteepRock to run a pharmacy and therapeutics (P&T) access analysis across three major US health systems: a large academic medical center, an integrated payer-provider system, and a community health system. The client brought its own prompts for the work. The analysis ran in Embedded Intelligence (EI), connected to the client's engagement records in OLMS and to OLA's licensed data.
Partway through, the analysis produced a payer mix for one of the largest treatment programs in the country. The figure looked fine. It was built from 231 pharmacy claims.
The licensed national dataset in the same session held 447,957 claims for the same drug codes over the same window. One of the biggest programs in the country accounting for 0.052 percent of national volume is not plausible, and EI said so. It compared the institutional figure to the national total, failed it, declined to present the payer mix, and explained why: the connected linkage data tied only 160 providers of any specialty to the institution, and only 16 in the relevant specialty. The math was fine. The provider roster underneath it was thin. The fix is a better roster, not a better prompt.
A stand-alone assistant would have presented that payer mix. It has no national total to check against, so it reports what it finds, and it reports it with the same confidence whether the number is right or wrong.

What else the session found
Four more findings, each from data the client's own analysis had not reached.
A conflict the client had not priced in. Open Payments records showed the client had paid the head of the academic center's clinical program roughly $27,000 in consulting and related fees from 2023 to 2025, most of it in 2025. Under a standard conflict-of-interest policy, the most influential clinical voice on this formulary decision may be the one person required to recuse from it.
An assumption that ran backwards. Trial registry data showed both academic systems had active trial relationships with a competitor developing a product in the same class, and neither had a confirmed trial relationship with the client. The client's earlier analysis assumed the reverse.
A quality measure that penalizes new products at launch. The relevant quality measures run off coded value sets that will not include a newly approved product on day one. Patients started on it will score as non-compliant, automatically, until those value sets update. That argues for timing protocol changes to a calendar-year boundary.
Two corrected premises. EI flagged two assumptions in the client's own analysis as unsupported by the connected data and showed the records behind each. The client is reviewing both.
It also declined to do several things it was asked to do. It would not fabricate committee rosters, would not invent base rates, would not assert trial participation it could not confirm, labeled its reconstructions as reconstructions, and stopped to ask when a prompt was ambiguous.
Why it worked
Three things, and the model is not one of them. EI runs Claude, the same model available in any enterprise seat. What changes the result is what the model is connected to, the rules it works under, and how the session is run.
The data. EI ran in the client's own private, access-controlled workspace. Everything in the table below was queried live during the session. The plausibility check that caught the bad number exists only because the national total was in the room to compare against.
Embedded Intelligence
Data source
Value @ $144/hr
Stand-alone assistant
Your engagement records in OLMS
Connected
No access
~0.5
Licensed national pharmacy claims
Connected
No access
Provider-to-institution linkage
Connected
No access
Open Payments records
Connected
No access
Clinical trial registries
Connected
No access
Certified treatment center list (about 8,700 centers)
Connected
No access
Your approved product and context documents
Connected
No access
~2.2
The rules. Every question in EI, whether it comes from the team's shared prompt library or is typed fresh, runs under written rules that SteepRock engineers, tests, and maintains: no guessing where data is missing, every figure labeled as confirmed or inferred, gaps declared, never filled, and every claim traced to a connected source, an attached file, or the user's own input. Record-level claims link back to the original entry in OLMS. Public information is cited.
The working method. The client's own prompts for this use case, four of them, were well written by any standard. SteepRock's team still rebuilt them into a ten-prompt library before the session, making fourteen corrections, writing down the expected result for each prompt, and grading every output against that standard. The rebuild is part of why the bad number was caught: the client's original sequence asked for the payer breakdown first and the data-quality check second, and a caveat delivered after a number does not undo the number. The rebuild made the capture check a precondition. EI's rules would have labeled the 231-claim figure as unreliable; the rebuilt prompt let EI withhold it altogether. That is the difference between a platform and a platform with a tested prompt library, and SteepRock delivers both. The client's original prompts, and what the rebuild changed in each, are in the appendix.
Why this matters now
Medical affairs teams at large manufacturers are moving generative AI from pilots into everyday work, and the regulatory expectation is moving with them. FDA's January 2025 draft guidance and the FDA and EMA's joint Guiding Principles of Good AI Practice in Drug Development, published in January 2026, expect AI whose outputs can be audited, and the EU AI Act's high-risk obligations, deferred by the Digital Omnibus adopted in July 2026, apply from December 2, 2027 for standalone systems and August 2, 2028 for AI embedded in regulated products [5]. The practical requirement is the same in every case: a team has to be able to show where a number came from and why it can be trusted. A screenshot of a chat does not meet that bar. A documented session, with expected results written down before execution and every output reviewed against them, does.\
What this means for your team
Medical and scientific teams get answers grounded in real records, every claim labeled as confirmed or inferred, and limitations stated next to the findings.
Account and access teams get the questions that decide a launch answered from licensed data a general assistant cannot reach: who decides, who pays, who influences, and where competitors already stand.
Compliance teams get a paper trail: expected results documented before execution, every output reviewed against them, and every claim traceable to its source, delivered as part of the engagement.
Bring us your prompts
SteepRock works exclusively with life sciences manufacturers in pharma, biotech, and medical devices, and EI's rules were written for that regulated environment. If your team already has a set of prompts it relies on, send us your ten best. We will rebuild them, write down the expected results, run them against your connected data in EI, and show you what changes. The client's identity in this case study is protected; if you would like to hear their experience directly, contact us and we will ask.
