Skip to content

OpenAI rolls out ChatGPT Health to U.S. users with a weaker AI model for free-tier accounts

· by Pondero Newsdesk

The short version

ChatGPT Health reaches U.S. users aged 18 and older on July 23, letting them connect Apple Health and medical records, but free accounts get GPT-5.5 Instant while paid subscribers get the stronger GPT-5.6 Sol for health analysis.

OpenAI rolls out ChatGPT Health to U.S. users with a weaker AI model for free-tier accounts

OpenAI began rolling out "Health in ChatGPT" to U.S. users aged 18 and older on July 23, 2026, six months after announcing the feature in January. Free-tier accounts get health analysis from GPT-5.5 Instant. Paying subscribers get GPT-5.6 Sol. That gap between the two models on medical benchmarks is substantial.

What the feature does

Users can connect Apple Health, medical records, and third-party wellness apps to ChatGPT to review lab results, prepare for doctor appointments, and analyze sleep or activity data, per The Decoder's coverage of the launch. OpenAI said it will not use connected health data for model training or advertising.

The feature originally launched with the Health section as a separate ChatGPT destination. OpenAI's own early testing found that more than 70 percent of users asked health questions outside the dedicated Health tab because switching was cumbersome. The company has since made the feature available in any conversation, with the separate Health section kept for managing data and reviewing past health chats.

Where the paid tier gap shows up

On OpenAI's internal HealthBench Professional benchmark, GPT-5.6 Sol scored 88.0 percent on completeness versus 53.2 percent for GPT-5.5 Instant, per The Decoder. The gap in health decision helpfulness ran 83.0 percent versus 50.8 percent. OpenAI said both models "beat doctors' answers" on the benchmark, with more than 260 physicians involved in developing the feature.

Those benchmark comparisons come with caveats. Physicians taking a written test may be under time pressure, working without patient records or colleague input, and may score lower than their clinical performance would suggest. OpenAI acknowledged in the announcement that ChatGPT can still make mistakes and cannot replace medical advice.

Independent benchmark data points in a different direction for diagnostic imaging. In the RadLE 2.0 radiology benchmark, none of the 16 AI models tested matched human radiologist performance. The central problem identified was overconfidence: the models returned incorrect findings with high certainty rather than acknowledging the limits of their analysis. Human radiologists were better at flagging when they had reached those limits.

Why operators and users should pay attention

The two-tier model structure raises a question that goes beyond pricing: whether stratifying clinical-quality AI by subscription status is the right design for a product that more than 300 million people use weekly for health questions, per OpenAI's own figure.

For AI-tool operators building health-adjacent workflows, the announcement signals that OpenAI is treating health as a platform surface, not a one-off feature. The data-connection layer (Apple Health, medical records, wellness apps) is an integration point, and operators who build workflows that touch health data will need to account for the model tier their end users are on.

The feature remains unavailable in the European Economic Area, Switzerland, and the United Kingdom. OpenAI cited no specific date for an EU launch. The exclusion reflects GDPR data-handling constraints and the likelihood that European regulators would classify the feature as high-risk under the EU AI Act, which could trigger conformity assessments and audit requirements before any rollout.

What to watch next

EU regulatory classification is the clearest outstanding question. If national authorities formally label ChatGPT Health as a high-risk AI system, OpenAI would face a defined compliance path before the feature can launch in Europe. Outside that track, independent replications of the HealthBench Professional benchmark using the public GPT-5.5 Instant and GPT-5.6 Sol APIs would give a clearer picture of whether the stated performance gaps hold in practice.

Sources