Clear, practical technology insights BSOD Code Lookup · Windows Error Code Lookup · Wi-Fi Troubleshooting · PC Troubleshooting Checklist

What a Flawed Apple Health Test Revealed About ChatGPT Health

A Washington Post experiment produced unstable cardiovascular grades from years of Apple Watch data. Here is what the test showed, what it did not prove, and how to use connected health data more safely.

Table of Contents

A Washington Post experiment found that ChatGPT Health gave unstable and clinically unhelpful cardiovascular grades after analyzing a reporter’s Apple Health history. The episode is a useful warning about asking a general AI system to convert wearable data into a diagnosis-like score. It does not establish that every ChatGPT Health response is wrong, and it was not a controlled clinical study.

Health in ChatGPT can connect supported medical records and Apple Health data for eligible users. OpenAI describes it as a tool for understanding information, spotting changes, and preparing for conversations with health professionals. It should not be used to diagnose a condition, rule out a disease, or decide whether urgent care is needed.

What happened in the Washington Post test

Technology columnist Geoffrey A. Fowler connected roughly a decade of Apple Watch and Apple Health data, including about 29 million steps and millions of heart-rate measurements. He then asked ChatGPT to assign a simple A-to-F grade to his cardiovascular health.

Apple Health's integration of ChatGPT Health has been criticized for serious misdiagnosis. Picture 1

According to Fowler’s January 2026 report, one response produced an alarming F. His physician did not agree with that assessment, and cardiologist Eric Topol also criticized the reasoning. Repeating or reframing the request produced substantially different grades, reportedly ranging from B to F.

The inconsistency matters because a health grade looks authoritative even when the scoring system is invented for the conversation. ChatGPT was not applying a validated cardiovascular-risk calculator with defined inputs and thresholds. It was being asked to compress years of mixed-quality data into a letter grade that medicine does not ordinarily use.

Why wearable data can be misread

Apple Watch and Apple Health can provide useful trends, but a long export is not the same as a complete medical record. Several limitations can distort an AI summary:

  • Measurements have context. Heart rate during sleep, exercise, illness, medication changes, and stress cannot be interpreted the same way.
  • Derived estimates are not diagnoses. Values such as estimated cardio fitness or heart-rate variability depend on device algorithms, sampling, fit, activity type, and available data.
  • Missing periods can bias a trend. A device that was not worn, replaced, or configured differently creates gaps that may look meaningful when they are not.
  • Risk requires information outside the watch. Age, symptoms, blood pressure, cholesterol, smoking, family history, examination findings, and validated tests can materially change a cardiovascular assessment.
  • A generated score may have no clinical basis. If the model invents a grading rubric, two plausible-sounding answers can disagree without either reflecting medical risk.

What ChatGPT Health currently does

OpenAI introduced ChatGPT Health in January 2026 and began a broader U.S. rollout in July. Its current product announcement says eligible U.S. users age 18 or older can access Health on web and iOS across Free, Go, Plus, and Pro plans. Availability can change by country, account, and platform.

Users can ask health questions without connecting a data source. When they choose to connect Apple Health or supported medical records, OpenAI says the service asks permission before using that information unless the user changes the setting to allow ongoing access. The company’s Health help page says connected Apple Health and medical-record data are read-only: ChatGPT cannot write information back to those sources.

OpenAI also says connected health information and conversations that use it are not used to train its foundation models or target advertising. Anyone considering the feature should still read the current Health privacy notice, review what is connected, and decide whether the benefit justifies sharing sensitive data.

Useful and unsafe ways to use the feature

Lower-risk supportDecisions that require a clinician
Summarize a trend to discuss at an appointmentDiagnose or exclude a disease
Draft questions about a lab result or wearable alertGrade cardiovascular health from raw wearable data
Organize a medication or symptom timeline for reviewStart, stop, or change prescribed treatment
Explain unfamiliar terms in plain languageDecide that new or severe symptoms are safe to ignore
Compare the wording of records while preserving uncertaintyChoose emergency, surgical, or insurance action from a chatbot answer alone

A productive prompt names the source and asks for uncertainty. For example: “Summarize the month-to-month resting-heart-rate trend, identify missing periods, and list questions I should ask my clinician. Do not diagnose or assign a health grade.” The result still needs checking against the original data.

How to check an AI health response

  1. Ask what evidence supports each conclusion. A response should distinguish measured values, estimates, and assumptions.
  2. Look for missing context. Verify dates, device changes, medications, illnesses, and the reason a measurement was taken.
  3. Request the method. If the answer gives a risk score or grade, ask whether it uses a named, validated clinical tool and whether all required inputs are present.
  4. Repeat cautiously. Large changes after a minor rewording signal that the conclusion is not robust; repeating until a reassuring answer appears is not validation.
  5. Bring the source data to a qualified professional. A clinician can decide whether the pattern is relevant and whether testing is warranted.

New chest pain, severe shortness of breath, fainting, stroke symptoms, or another possible emergency should be handled through local emergency services, not an AI chat.

The Fowler test exposed a real failure mode: an unsupported letter grade appeared more precise than the underlying evidence. The safer role for ChatGPT Health is to help users organize information and prepare better questions while keeping clinical judgment with qualified health professionals. For broader context on phone-based assistants, see our guide to useful AI apps for iPhone and our overview of Apple Watch health features.

Discussion

Reader Comments 0

Sign in with email or Google to join the discussion.