Most of us answer clinical questions with some mix of three things: a curated reference database like UpToDate or DynaMed, an AI evidence tool like OpenEvidence or UpToDate Expert AI, and a general-purpose model like ChatGPT or Gemini. Choosing among point of care clinical tools in 2026 depends less on which one wins a benchmark and more on the kind of question you are asking, and on how much verification the answer needs before it reaches a patient. Two independent evaluations published this summer reached opposite conclusions about the same tools, which tells you something important about how to read any vendor's accuracy claim.
This guide covers what the current evidence actually shows, where each category of tool breaks down, the regulatory assumption you are standing on every time you use one, and how to build a workflow you could defend in a chart review.
The problem was never access to information
Here is the finding that should anchor how you think about this. In a systematic review of 72 studies, clinicians raised an average of 0.57 questions per patient seen. They pursued only 51 percent of those questions. When they did pursue one, they found an answer 78 percent of the time [1].
Read that again. We are good at finding answers. We just do not go looking for half of them.
The two barriers the authors identified were lack of time and doubt that a useful answer existed at all [1]. Neither is a knowledge problem. Both are workflow problems, and they have proven remarkably stubborn. That review pooled data spanning decades of expanding online access, and the pattern barely moved.
The time pressure has not improved since. Primary care physicians in one three-year event-log study spent 355 minutes, roughly 5.9 hours, of an 11.4-hour workday inside the EHR, with 86 minutes of that falling after clinic hours [2]. Clerical work accounted for 44 percent of EHR time and inbox management for another 24 percent [2]. The margin left for looking something up is thin, and it competes directly with going home.
Meanwhile the corpus keeps growing. PubMed logged its 20 millionth citation in July 2010 and its 35 millionth in December 2022. It now holds more than 41 million [3]. The literature roughly doubled in a dozen years while your available minutes did not.
So the honest framing of the 2026 tool question is not "which product knows the most medicine." It is "which product gets me to a defensible answer inside the time I actually have, and how much checking does that answer still require."
The four categories, and what each one is actually for
Vendors blur these lines deliberately. Keeping them separate makes your choices clearer.
Curated reference databases. UpToDate, DynaMed, BMJ Best Practice, ClinicalKey. Human authors write and peer review topic monographs, editors update them, and you navigate to the relevant section. UpToDate alone lists more than 13,000 clinical topics across 25 specialties, maintained by a stated 7,600-plus physician authors, editors, and reviewers [5]. The output is slow to read and slow to change, which is both the weakness and the entire point.
AI evidence-synthesis tools. OpenEvidence, UpToDate Expert AI, Dyna AI, ClinicalKey AI, Doximity Ask. You ask a question in plain language and get a synthesized answer with inline citations. If you want the product-by-product detail, we break these down separately in UpToDate versus OpenEvidence versus the new AI tools. These are the fastest option by a wide margin, and they are now the most-used AI application in American medicine. In the AMA's 2026 survey, "summaries of medical research and standards of care" was the single most common AI use case, incorporated by 39 percent of physicians, up 26 percentage points from 2024 [4].
Calculators, drug references, and guideline apps. MDCalc, Lexidrug, Epocrates, Sanford Guide, specialty society apps. Narrow scope, deterministic output, essentially no hallucination risk for the calculation itself. Still the correct tool for a Wells score or a renal dose.
General-purpose frontier models. ChatGPT, Claude, Gemini. Strong reasoning, no clinical governance by default, and no business associate agreement unless your institution has one. This is the category that carries the most risk and, according to the 2026 benchmark data, sometimes the most capability.
What the 2026 evidence actually says, and why two studies disagreed
This is where most tool roundups stop being useful, so it is worth going slowly.
On June 12, 2026, Nature Medicine published a brief communication from a team at NYU Langone comparing two specialized clinical tools, OpenEvidence and UpToDate Expert AI, against three frontier models: GPT-5.2, Gemini 3.1 Pro, and Claude Opus 4.6 [10]. Three stages. Five hundred MedQA licensing-style questions, 500 HealthBench items measuring alignment with clinicians, and a real clinical queries benchmark built from 100 de-identified questions physicians had actually typed into NYU's HIPAA-compliant model instance, scored by 12 blinded clinicians producing 1,800 annotations [10].
The frontier models won all three stages. On MedQA, Gemini scored 97.4 percent against 89.6 percent for OpenEvidence and 88.4 percent for UpToDate Expert AI. On HealthBench, GPT scored 88.0 against 62.6 and 61.3 for the two clinical tools. On the real-query benchmark, the specialized tools performed no better than Google's automatic Search AI Overview [11].
Two weeks later, a preprint went up on arXiv with the opposite result. The Real-POCQi evaluation used 620 genuine point-of-care queries submitted to OpenEvidence across 30 specialties, plus 187 HealthBench items, and recruited 149 practicing physicians across 36 states to make blinded head-to-head comparisons, with each grader matched to the specialty of the question [12]. OpenEvidence won on all five dimensions tested, with win margins between 25 and 39 percentage points over Claude Opus 4.8, Gemini 3.1 Pro, and GPT-5.5 [13].
Neither study is fraudulent. They tested different questions, drawn from different user populations, judged by different people, on different model versions, using different scoring methods. Of course they diverged.
A few things are worth naming plainly. The contamination concern OpenEvidence raised, that MedQA and HealthBench may sit in the frontier models' training data, is legitimate, and the NYU authors conceded it in the paper, which is why they designated the real-query benchmark as their primary evidence [11]. It is also true that a benchmark built entirely from one vendor's own user base will favor that vendor, because the questions people bring to a tool are shaped by what they have learned it does well. And it is true that OpenEvidence responded to the Nature Medicine paper by demanding a retraction and a public apology rather than submitting a rebuttal through the journal's Matters Arising process [11].
The practical takeaway is not that one tool won. It is that in 2026 there is still no independent, prospective, outcomes-based evidence that any of these products improves patient care. Notably, none of the models in the NYU study produced more harmful content or more hallucinations than the others [11]. The differences were in completeness and communication, not safety.
Which brings up the older evidence that vendors still lean on. Wolters Kluwer's marketing describes UpToDate as the only clinical decision support resource associated with improved patient outcomes. The study behind that claim is real: a 2012 retrospective analysis of Medicare data found that patients at hospitals using UpToDate had shorter stays, 5.6 days against 5.7, and lower risk-adjusted mortality for three of six conditions, with reductions between 0.1 and 0.6 percent [17]. Read the effect sizes rather than the headline. The authors themselves characterized the association as very small but consistent, concentrated in smaller non-teaching hospitals, and it is an association, not a causal finding.
Similarly, the Mayo Clinic study OpenEvidence points to for validation examined five retrospective patient cases across five chronic conditions. It rated the tool highly for clarity, relevance, and evidence support, and found the impact on clinical decision making to be minimal, with the authors calling for prospective trials [18]. That is a reasonable pilot. It is not a reason to change your practice.
The regulatory line you are standing on every time you use these tools
Almost no physician I talk to knows this, and it changes how you should read a citation list.
None of these tools are FDA-cleared medical devices. They operate inside a carve-out created by the 21st Century Cures Act, which excludes certain clinical decision support software from the device definition if it meets four criteria. The fourth is the one that matters to you: the software must be intended to let the health care professional independently review the basis for its recommendations, so that you are not relying primarily on the recommendation to make the decision.
That is the deal. The citation trail is not a courtesy feature. It is the legal premise on which the entire category sits outside device regulation.
The FDA rewrote its interpretation of these criteria in January 2026, issuing revised final guidance on January 6 and re-issuing it on January 29, superseding the September 2022 version, then holding an industry town hall on March 11 [14][15]. The changes matter for developers more than for clinicians, but two are worth knowing. The agency loosened its position that non-device software must offer multiple options, saying it intends to exercise enforcement discretion when only one option is clinically appropriate. And it moved the time-critical decision-making limitation into criterion four, on the reasoning that you cannot meaningfully review the basis for a recommendation when the clock is running [14].
Read that last one against your own habits. If you are using a fast AI answer precisely because there is no time, you are using it in the exact scenario the FDA flags as incompatible with independent review.
The safety community reached a related conclusion from a different direction. ECRI put misuse of AI chatbots in health care at the top of its 2026 list of health technology hazards, its first entry that is not a medical device at all [16]. Worth noting the scope carefully, because coverage of this report has been sloppy: ECRI tested commonly available general-purpose chatbots, and specifically excluded purpose-built clinical applications such as OpenEvidence and ChatGPT Health from that testing [19]. The finding is about pasting a clinical question into a consumer chatbot, not about the specialized tools.
The question nobody in these comparisons asks: who is paying?
OpenEvidence is free to verified United States clinicians, and it is free because pharmaceutical and device manufacturers pay to advertise inside it. As of early 2026 the company reported roughly 757,000 registered clinicians and more than 20 million clinical consultations per month, with a $12 billion valuation after a January 2026 Series D [8][9]. The advertising commands CPM rates far above ordinary digital media, for an obvious reason: the ad is served at the moment a physician is deciding what to do [9].
I am not claiming this corrupts the evidence. I have seen no data showing it does. But it is a structural conflict that belongs in your assessment, and it is missing from essentially every comparison article currently ranking for these keywords, most of which are published by companies selling a competing tool. When you read that OpenEvidence is used daily by more than 40 percent of American physicians, note that this is a company figure, not an independently audited one [8].
The subscription model has its own distortion. Wolters Kluwer does not publish a list price for personal UpToDate subscriptions on its main product pages, selling instead through its own webstore and through discounted society programs with the AAFP, AAPA, and ACOFP [6]. Third-party sources report 2026 individual prices ranging from roughly $549 to $699 per year depending on tier, and they disagree with each other, so confirm the current figure in the webstore before you budget for it [UNVERIFIED PRICE RANGE, AUTHOR TO CONFIRM IN WEBSTORE]. Expert AI, the generative layer, sits in the higher Pro Plus tier [6]. Before you pay anything, check whether your hospital credentialing, state medical society, or local library already provides access.
Doximity occupies a third position. Ask and Doximity GPT are free to verified United States clinicians, and Doximity's own documentation states that PHI may be included in prompts under HIPAA-compliant protocols with encryption in transit and at rest [7]. For a physician without institutional coverage, that combination is unusual and underrated.
How to choose: match the tool to the clinical job
Skip the composite scores. Ask what kind of question you have.
Is it a calculation, a dose, or an interaction? Use a calculator or drug reference. Do not ask a language model to compute a Wells score. The deterministic tool cannot be confidently wrong in the way a model can.
Is it a bounded factual question where you will act on the answer today? An AI evidence tool is the right call, on one condition: you open at least one primary citation before the answer changes management. If you will not do that, use a curated reference instead. The FDA carve-out and your own liability both assume you looked.
Is it a management question on a patient who is complicated or unfamiliar to you? Curated reference, every time. You want the disagreements, the caveats, and the named human who wrote it. AI synthesis tends to flatten controversy into a clean recommendation, and clean recommendations on complex patients are exactly where you get hurt.
Is it a reasoning problem rather than a lookup problem? Something like organizing a differential on an atypical presentation. Frontier models are genuinely strong here and the NYU data supports that [10]. Two rules: no PHI without a BAA, and treat everything it produces as a hypothesis to check, not an answer. How you frame the question changes the output more than most physicians expect, which is the subject of our guide to getting reliable clinical answers from AI.
Is it a question you have now asked three times this month? That is a knowledge gap, not a lookup. Read the actual guideline or review once, properly. Tools that give you a fast answer every time can quietly prevent you from ever learning the thing, which is precisely the concern 88 percent of physicians reported in the AMA survey, rising to 35 percent worry about personal skill loss among those under 10 years in practice [4].
If most of your questions land in category two, that is the specific job Physicians' Copilot was built for. Try Physicians' Copilot for point-of-care answers and see whether the citation trail holds up on the questions you actually ask, which is the only test that matters.
Guardrails worth writing down
PHI does not go into a tool without a BAA. Physicians in the AMA survey were substantially more worried about privacy with non-institutional tools, 71 percent, than with institutional ones, 42 percent [4]. That instinct is correct. De-identify or use a covered tool.
Verify anything that changes management. Not everything. That standard is unworkable and nobody meets it. The workable version is a threshold: if the answer changes what you do, open the citation.
Check the date on the underlying evidence, not the interface. A confidently written 2026 answer can be built on a 2019 guideline. Both curated and AI tools have this failure mode.
Watch for the fluency trap. A well-written wrong answer is more dangerous than a poorly written one, because it does not trigger the skepticism a garbled answer would. When the NYU authors found OpenEvidence scoring lowest on clarity, they attributed it to communication rather than knowledge [11]. The inverse is the risk you should worry about.
Know what your institution has actually approved. The AMA found 73 percent of physicians have received some AI training, but only 11 percent of those describe it as substantial, and 92 percent want more [4]. If your organization has no governance policy, your personal use is the policy. Scribes, liability, and institutional governance sit outside the scope of this guide and are covered in our practical guide to AI in clinical medicine.
No AI tool replaces clinical judgment. Verify all outputs against primary sources and your own assessment.
Key takeaways
Roughly half the clinical questions you raise never get pursued, and the barriers are time and doubt, not access. Pick tools that shrink time to a defensible answer [1].
Two 2026 evaluations of the same tools reached opposite conclusions because they tested different questions with different graders. Treat any single accuracy claim, from any vendor, as provisional [10][12].
These products are not FDA-cleared devices. They sit in a Cures Act carve-out that assumes you independently review the basis for the recommendation, which the FDA reinforced in its January 2026 revised guidance [14].
Match the tool to the job: calculators for computation, AI synthesis for bounded questions, curated references for complex management, frontier models for reasoning with no PHI.
Ask who is paying. Free clinical AI is generally funded by pharmaceutical advertising served at the moment of decision, which is a conflict worth holding in mind even absent evidence of harm [9].
Where to go from here
The category is moving faster than the evidence, and it will keep doing that for a while. What will not change is the standard you are held to, which is that you reviewed the basis for the recommendation before you acted on it. Build the workflow around that and the tool choice gets much simpler.
Try Physicians' Copilot for point-of-care answers on your next shift, and judge it the way you should judge all of them: pick three questions you already know the answer to, and see whether the citations hold.
REFERENCES
All URLs accessed August 22, 2026.
Del Fiol G, Workman TE, Gorman PN. "Clinical Questions Raised by Clinicians at the Point of Care: A Systematic Review." JAMA Internal Medicine, May 2014;174(5):710-718. https://pubmed.ncbi.nlm.nih.gov/24663331/
Arndt BG, Beasley JW, Watkinson MD, Temte JL, Tuan WJ, Sinsky CA, Gilchrist VJ. "Tethered to the EHR: Primary Care Physician Workload Assessment Using EHR Event Log Data and Time-Motion Observations." Annals of Family Medicine, September 2017;15(5):419-426. https://www.annfammed.org/content/15/5/419.short
National Library of Medicine. "PubMed Milestone: 35 Millionth Journal Citation Added," NLM Technical Bulletin, December 14, 2022; current count per NCBI database description. https://pubmed.ncbi.nlm.nih.gov/
American Medical Association, Center for Digital Health and AI. "2026 Physician Survey on Augmented Intelligence." March 2026. https://www.ama-assn.org/system/files/physician-ai-sentiment-report.pdf
American Academy of Family Physicians. "UpToDate Subscription Discounts" member benefit page. https://www.aafp.org/membership/benefits/advantage/uptodate
Wolters Kluwer. "UpToDate Pro Plus" product page and "UpToDate and Lexidrug Subscription Options." https://www.wolterskluwer.com/en/solutions/uptodate/pro/uptodate/uptodate-pro-plus
Doximity Support Center. "Doximity GPT FAQs." https://support.doximity.com/hc/en-us/articles/41850759500051-Doximity-GPT-FAQs
CNBC. "OpenEvidence, the 'ChatGPT for doctors,' doubles valuation to $12 billion." January 21, 2026. https://www.cnbc.com/2026/01/21/openevidence-chatgpt-for-doctors-doubles-valuation-to-12-billion.html
MedCity News. "Thunderstruck By OpenEvidence's $12B Valuation? Don't Be." February 12, 2026. https://medcitynews.com/2026/02/openevidence-healthcare-valuation/
Vishwanath K, Oermann EK, et al. "General-purpose large language models outperform specialized clinical AI tools on medical benchmarks." Nature Medicine, June 12, 2026. https://www.nature.com/articles/s41591-026-04431-5
Bruce G. "ChatGPT, Gemini, Claude outperform clinical AI tools: Study." Becker's Hospital Review, June 16, 2026. https://www.beckershospitalreview.com/healthcare-information-technology/ai/chatgpt-gemini-claude-beat-clinical-ai-tools-study/
Ouyang J, Vossler P, Feng J, et al. "Expert Evaluation of Clinical AI Tools on Real Point-of-Care Clinical Queries." arXiv preprint 2606.28960, June 27, 2026. https://arxiv.org/abs/2606.28960
Becker's Hospital Review. "OpenEvidence beats Claude, Gemini, GPT-5.5 in new physician-led study." July 2026. https://www.beckershospitalreview.com/healthcare-information-technology/ai/openevidence-beats-claude-gemini-gpt-5-5-in-new-physician-led-study/
U.S. Food and Drug Administration. "Clinical Decision Support Software: Guidance for Industry and Food and Drug Administration Staff." Issued January 6, 2026; re-issued January 29, 2026. https://www.fda.gov/media/109618/download
U.S. Food and Drug Administration. "Town Hall: Clinical Decision Support Software, Final Guidance." March 11, 2026. https://www.fda.gov/medical-devices/medical-devices-news-and-events/town-hall-clinical-decision-support-software-final-guidance-03112026
ECRI. "Misuse of AI chatbots tops annual list of health technology hazards" (Top 10 Health Technology Hazards for 2026). January 21, 2026. https://home.ecri.org/blogs/ecri-news/misuse-of-ai-chatbots-tops-annual-list-of-health-technology-hazards
Isaac T, Zheng J, Jha A. "Use of UpToDate and outcomes in US hospitals." Journal of Hospital Medicine, February 2012. DOI 10.1002/jhm.944. https://shmpublications.onlinelibrary.wiley.com/doi/10.1002/jhm.944
Hurt RT, Stephenson CR, Gilman EA, Aakre CA, Croghan IT, Mundi MS, Ghosh K, Edakkanambeth Varayil J. "The Use of an Artificial Intelligence Platform OpenEvidence to Augment Clinical Decision-Making for Primary Care Physicians." Journal of Primary Care & Community Health, April 16, 2025;16. https://journals.sagepub.com/doi/10.1177/21501319251332215
Betsy Lehman Center for Patient Safety. Coverage of ECRI Top 10 Health Technology Hazards 2026 webinar, including tested-product scope. February 2026. https://betsylehmancenterma.gov/news/normalization-of-ai-chatbots-presents-risks-and-benefits-in-health-care-or-ai-chatbots-top-ecris-list-of-technology-hazards-for-2026
Deliberately excluded, and worth saying so: Wolters Kluwer told Becker's that UpToDate Expert AI returned clinically aligned information for 99.9 percent of assessed criteria across 1,669 queries. No methodology has been published for that evaluation, so it is not cited in the article. The same rule was applied to OpenEvidence's internal accuracy claims.
FAQs
The published evidence does not settle this. The June 2026 Nature Medicine study found OpenEvidence at 89.6 percent and UpToDate Expert AI at 88.4 percent on MedQA questions, with both trailing frontier general-purpose models, and both performing no better than Google's AI Overview on real physician queries [10][11]. A subsequent preprint using OpenEvidence's own user queries found the opposite ordering against frontier models [12]. They answer different questions. For complex management I still reach for a curated reference; for a fast bounded question I use AI synthesis and open a citation.