Jump to

    Research

    MedPAIR: Measuring Whether Physicians and AI Agree on What Matters in Medical QA

    Featuring Yuexing Hao, researcher at Microsoft and an MIT EECS postdoctoral associate.

    October 2, 2026
    •
    27 min read
    •

    Jump to

      At our latest Snorkel AI Reading Group, Yuexing Hao, co-founder and COO of MatrAIx and previously a postdoc in MIT’s Healthy ML Group, presented MedPAIR, a dataset that compares which parts of a clinical case physicians and LLMs treat as relevant. Physician trainees and board-certified physicians labeled every sentence in 933 medical exam questions, and six LLMs labeled the same sentences. The question it asks: when a model answers a clinical question correctly, is it reasoning from the same evidence a physician would use? MedPAIR was supported by Snorkel’s Open Benchmarks Grants program.


      Lightly edited for readability.

      Introduction

      Moderator: Yuexing Hao is co-founder and COO of MatrAIx, building a new evaluation paradigm for AI agents. Previously, she was a postdoc at MIT’s Healthy ML Group and holds a PhD from Cornell University. Her research focuses on AI for healthcare and human-computer interaction.

      This work is increasingly important as discussion grows around the future of human work and humanity in the age of AI, particularly in medicine. Snorkel is happy to have supported Yuexing’s work on MedPAIR through the Open Benchmarks Grants program. Please visit us if you would like to learn more. She will also present this work at the upcoming Frontier Data Summit on October 8.

      After the presentation, we’ll open the floor for Q&A and then break out for further discussion. Please join me in welcoming Yuexing.

      Presentation

      Yuexing Hao: Hello, everyone. Thank you for the warm welcome. I’m currently in Mountain View, so San Francisco is fortunately not too far away. It’s great to see so many people, including some familiar faces.

      Today I’m going to present my AI for Healthcare work, one of the most important projects from my PhD. I received my PhD from Cornell University this January. During the last two years of my PhD, I was at MIT EECS, mentored by Professor Marzyeh Ghassemi. This project became one of the central works of my thesis, and I’m proud to share it.

      The project is called MedPAIR, which stands for Measuring Physician–AI Relevance Alignment. We’re trying to measure how closely physicians and AI systems align in their reasoning about relevance, particularly in medical question answering.

      I want to acknowledge my collaborators at MIT, including Kumail Alhamoud, Hyewon Jeong, Haoran Zhang, Isha Puri, and Grace Yan, as well as Professor Philip Torr at Oxford, Mike Schaekermann, my former internship supervisor at Google Research, Saleh Kalantari, my Cornell advisor, Ariel D. Stern of HPI in Germany, and Professor Marzyeh Ghassemi at MIT. This is joint work with all of these collaborators.

      We’re also excited to present a poster at Snorkel’s Frontier Data Summit on October 8. Snorkel is also supporting our workshop at COLM on October 9. The workshop is called DAIH: LLM/VLM Deployment Opportunities and Risks in Healthcare. There will be exciting work on how people are deploying large language models in healthcare.

      The question: how do physicians and AI use context?

      When we started thinking about MedPAIR in 2024, one of the questions we kept asking was: Large language models are good, but are they really deployable? Why should we trust AI? Can these models be comparable to physicians? Those questions have become even more important as more companies work on medical large language models.

      There is already a lot of related research. For example, work in Nature has studied conversational diagnostic AI. Researchers at Stanford’s Human-Centered AI Institute have asked whether AI can improve diagnostic accuracy. Other work explores how to teach language models to think like clinicians, including research on sequential diagnosis with language models. Many major technology companies and frontier labs are working to make language models more trustworthy and more similar to physicians or physician trainees.

      We wanted to understand how to evaluate the differences between physicians and language models in clinical decision-making. Before explaining our work, I’ll briefly review what the field has found. A 2024 Nature Medicine paper reported a surprising result. You might expect that adding more patient information, such as laboratory tests, imaging, and physical examination results, would improve clinical decisions. More context should make predictions more accurate. But the results did not always follow that pattern. Adding information could sometimes lower diagnostic accuracy.

      That led us to a more fundamental question: When physicians and language models interpret the same patient context, do they focus on the same information?

      When you visit a clinician, they ask about your demographic information, symptoms, medications, and family history. How do physicians interpret that information? How do language models interpret it? Are their judgments alike?

      We know that language models can achieve performance similar to physicians on some clinical decision-making tasks. But do they use similar reasoning processes? Does a final answer reflect the model’s reasoning, or did it simply recall something from its training data? We wanted to ask two questions: Are physicians and models alike in what they consider relevant? And, if so, how are they alike?

      Study design

      We use a multiple-choice question-answering setup similar to the USMLE. A case includes a patient profile, a question, and several possible answers. In one example, a profile contains six sentences describing a previously healthy 40-year-old man with a three-month history of a slowly enlarging right breast mass and associated pain. The question asks for a diagnosis.

      Before studying chain-of-thought reasoning, we focused on the input context. We asked two types of judges, physicians and large language models, to label each sentence in the patient profile as highly relevant, low relevance, or irrelevant to answering the question.

      For example, a language model might label sentences one, three, and five as highly relevant, and sentences two, four, and six as irrelevant. Physicians might label sentences one, two, and five as highly relevant. Sentence three then becomes an interesting point of disagreement.

      Our study has two rounds. First, we collect sentence-level relevance annotations. Then we ask whether the sentences identified as relevant are actually necessary for making the clinical decision. If we keep only the relevant sentences and remove the rest, can a physician or model still answer correctly?

      For the human annotations, we worked with Snorkel AI to recruit 52 physician trainees and seven board-certified physicians. Annotators first had to answer the question correctly. If they did not answer correctly, we did not use their sentence labels. Those who answered correctly labeled every sentence as highly relevant, low relevance, or irrelevant. We aggregated the human annotations using majority vote.

      We also asked six language models to annotate the cases. We used both open- and closed-source models and did not fine-tune them. For each model, we collected two types of relevance judgments: self-reported judgments produced from a prompt, and context scores indicating how much weight the model assigned to each sentence when making its prediction (computed with ContextCite for the open-source models).

      The context score gives a number rather than one of the three human labels. To make the comparison fair, we used the same sentence budget for each question. We selected the top sentences according to the physician trainees’ majority-vote annotations, then selected the same number of top sentences according to each model’s context scores.

      The resulting MedPAIR dataset combines the original questions from public benchmarks with physician-trainee annotations, model context scores, and model self-reported relevance labels. The dataset is available on Hugging Face.

      We also define a spurious rate. The question is whether removing sentences labeled low relevance or irrelevant changes a model’s answer. We want to understand whether focusing on highly relevant information helps models make predictions. I’ll explain the formula in the results.

      Hypotheses and discussion with the audience

      Yuexing Hao: I’d like to try a few hypotheses with the audience. First, are highly relevant sentences consistently longer and more uniform in structure?

      Audience member: No.

      Yuexing Hao: Right. They are not necessarily longer, and they can be structurally complex. By uniform, I mean whether they contain surprising findings.

      Second, do humans and language models disagree about which information is relevant?

      Audience: Yes.

      Yuexing Hao: We’ll see how much. Third, can human relevance judgments improve language-model performance? If we remove information that physician trainees label as irrelevant, do you think performance improves?

      Audience: [Mixed responses.]

      Yuexing Hao: That’s exactly what we wanted to find out. A related question is whether a model can improve its own performance by using the sentences it identified as relevant.

      Audience member: Could results differ across models, depending on each model’s capabilities?

      Yuexing Hao: Exactly. Not every language model is capable of medical reasoning, and model capabilities may affect the results. The model’s cost and whether it was trained specifically for medicine may also matter.

      Audience member: Did you restructure the sentences? For example, reverse their order or convert them into YAML?

      Yuexing Hao: That was one of our first questions. We changed sentence order to test whether the perturbation affected the results. We did not see much difference in downstream performance, so we moved away from that research direction. Changing the order could still matter in some cases, especially when the timeline is important. If a patient’s symptoms occur after taking a medication, changing that sequence could change the narrative. But in our tests, sentence-order changes did not make much difference overall.

      Audience member: Did you also shuffle the answer choices? Research has shown that shuffling USMLE choices can confuse models.

      Yuexing Hao: Yes, changing the options can matter. Some datasets, such as MedXpertQA, have ten options, so the ordering could affect downstream performance.

      Audience member: I have a healthcare background and know clinicians who feel pressure as more healthcare companies adopt AI. If AI is being used, clinicians may be expected to do more or prove more value. Is it a good idea to keep a human in the loop for every task? In triage or scoring, for example, we don’t want clinicians to review everything. They need to focus on the tasks that require their attention.

      Yuexing Hao: That’s why we worked with Snorkel AI and included humans in the loop in this study. We wanted to understand how automated clinical decisions compare with decisions made by people. That is the fundamental question we’re starting with.

      I agree that keeping a human in the loop for every task can defeat the purpose of automation. In a high-pressure setting like an emergency room, triage or scoring could help clinicians decide which tasks to handle first. But healthcare is not only about accuracy. It is a whole service. Clinicians do more than make diagnoses or recommend treatments, so we need to think about responsible and ethical AI use. In some areas, language models may work well as decision-support tools while people remain responsible for care.

      My first paper in AI for healthcare was in human-computer interaction. Starting in 2023, I spoke with clinicians about how AI could support their decisions. At the time, many physicians did not want to see AI as a competitor. Now, with tools such as OpenEvidence and the growing popularity of medical language models, people are seeing this trend. But hospitals are still here, and people still need to go to them. My professor, Marzyeh Ghassemi, has joked that if you have to wait eight hours in an emergency room, you might ask the best medical AI your question instead. Even so, most people would still want to go to the hospital. Healthcare has a different nature, and our work is trying to address some of its critical challenges.

      Results

      Yuexing Hao: Let’s move from the hypotheses to the results. I spent almost a year working on them, so I’ll summarize them here.

      First, we measured agreement between humans and language models on information relevance. The heat map includes open models such as Qwen, Llama, and MedGemma, as well as closed models such as GPT-4o and GPT-5. We compare human labels with models’ self-reported judgments and context scores.

      Human and model judgments agree on roughly 50 to 60 percent of the sentences. That means there is substantial disagreement about which parts of the input are highly relevant. Models also disagree with one another. For example, agreement between some models is quite low, while Qwen 72B and Llama 70B show close to 80 percent agreement on context scores. The model’s capabilities and training may affect its judgments, and model developers do not always disclose the details needed to explain those differences.

      A case study shows how these disagreements appear at the sentence level. We compare what physicians consider highly relevant with what different models identify as irrelevant or low relevance.

      Next, we asked whether human relevance judgments improve model performance. In the chart, the black line represents the original accuracy and the dotted lines show the best performance. Filtering out sentences that physician trainees identified as irrelevant can substantially improve performance for some models, but not all of them. For example, one MedGemma result became worse after removing sentences selected by the human annotators. Model relevance judgments do not always align with human judgments. Using a model’s own relevance labels also did not consistently improve downstream performance. In some cases, simply making the context shorter and more focused helped.

      Our third question was whether a model could make the same prediction using only sentences labeled highly relevant. We call this the sensitivity of model predictions to sentence selection, and it relates to the spurious rate I mentioned earlier.

      Across 933 questions curated for MedPAIR, the relevant sentences generally preserve higher accuracy than the irrelevant sentences. But removing context can still change answers. For Qwen 72B, 46.3 percent of questions that were originally answered correctly became incorrect when only the highly relevant sentences were retained. For Llama 70B, that figure was 14.9 percent. GPT-5 appeared less affected by the perturbation, though we are also considering whether benchmark questions may have appeared in its training data. The random-sentence condition had surprisingly little effect.

      Another chart shows how often each sentence-selection strategy produces the best performance across the four datasets used in MedPAIR: MMLU, JAMA, MedXpertQA, and MedBullets. We compare relevant-only, irrelevant-only, and random sentence sets.

      To summarize, MedPAIR is a benchmark that compares relevance annotations from clinical professional labelers with relevance estimates from language models. We want to understand how models identify relevant information in clinical cases and how closely those judgments match what physician trainees find relevant.

      Next steps for MedPAIR

      Yuexing Hao: I’ve graduated and shifted somewhat from healthcare to broader domains, but the MedPAIR work is ongoing at MIT. We’re thinking about how to expand it.

      One direction is multimodal evaluation. The examples I showed use text, but medical cases can also include images, figures, and pathology reports. We need to understand which information models identify as relevant in those settings. It is relatively straightforward to label relevant sentences, but much harder to identify the relevant areas of an image.

      With support from Snorkel’s Open Benchmarks Grants program, we’re looking to expand along two dimensions. On the horizontal axis, we want to increase dataset size. We currently have 933 questions with irrelevant sentences in the patient profiles. We also want to add multimodal and multilingual benchmarks. On the vertical axis, we want to improve input-relevance accuracy, including by studying fine-tuned models. We’re also aiming to add more data types, including images, video, and audio, more questions, and ratings from board-certified physicians. Snorkel’s support is important to this work.

      MatrAIx and simulation-based evaluation

      Yuexing Hao: I’ll also briefly talk about what I’m working on now. MedPAIR is a long-running project, and people at MIT are continuing to explore it. My current work focuses on evaluation.

      Language models can perform well, but judging them is difficult. Frontier labs may launch a new agent, but online evaluations with human analysts are time-consuming and expensive. It can take one or two months to get the first batch of results. At the same time, many static offline benchmarks have become saturated. Models score well on benchmarks that were challenging just a few months ago, while real users may still be dissatisfied with AI agents.

      MatrAIx is trying to bridge the gap between online and offline evaluation. We started by curating a persona dataset with 8.3 billion personas. Some are synthetic, and some are based on real-world data. We use a directed, attribute-based graph to represent how the personas are generated across 1,290 dimensions.

      We built a simulation-based evaluation infrastructure that can support different stages in the lifecycle of an AI product. First, a team may use surveys to understand which capabilities it wants to target and compare different concepts or variants. Then it can evaluate a model through multi-turn interactions, checking for hallucinations, red-team issues, quality, latency, or verbosity. Finally, once the agent is part of a product such as a website or mobile app, the team can evaluate task completion, screenshots, and failure cases.

      The goal is to fill the gap between online evaluation and static offline benchmarks with simulation-based evaluation. We call the testing environment a sandbox.

      For example, a Coca-Cola team might ask whether current customers would keep buying after a two-dollar price increase. Relevant persona agents take part in a simulated survey and provide answers and rationales. In one example, 61 percent of one million persona agents said they would still buy the product after the price increase, compared with a target of 55 percent. The agents are synthetic, but the interaction histories, or telemetry, are real records of how they interacted with the survey.

      Audience Q&A: persona representativeness

      Audience member: You have a large number of simulated personas and concrete results. Does the distribution of personas reflect the true distribution of people? If not, the numbers may not transfer to real-world use cases.

      Yuexing Hao: That’s a question we get often. Before publishing the paper, we checked whether the synthetic users matched real-world distributions, including gender, ethnicity, and race. There are also subjective attributes, such as people’s preferences for coding agents. For those, we model correlations and dependencies among persona attributes. We do not want a persona to have combinations that do not make sense, such as being 18 years old with 17 years of work experience.

      After we released the paper, we received press attention, including two articles in Forbes. We are also speaking with individual contributors and small companies that are deploying MatrAIx in their evaluation pipelines. We’re collaborating with Stanford, new research labs, frontier labs, and larger technology companies to integrate the evaluation infrastructure into their pipelines.

      We currently have more than 1,000 tasks across over 25 domains, including e-commerce, software, finance, and healthcare. We evaluate four types of environments: surveys, AI bots, websites, and apps. We measure 91.5 percent overall adherence between the simulated results and human judgments. Adherence is higher in software testing and lower in subjective areas such as humor, politeness, storytelling, and jargon. Survey, chat, and web environments show more than 90 percent alignment, while OS apps are lower, at around 83 percent.

      Some tasks are inherently time-sensitive. For example, asking whether someone would buy a stock today or tomorrow could produce different answers depending on the market. That is part of what we need to capture in human validation.

      The research project began as an open-source research community. We’re grateful to have 18 advisors with experience in user simulation, evaluation infrastructure, and AI infrastructure, including leaders from Microsoft, Liquid AI, Meta, Google DeepMind, and Oracle, as well as professors from MIT, Harvard, Stanford, and Oxford.

      Our slogan is “Explore the future today and simulate before reality.” Right now, we focus on AI agent simulation and evaluation. Thank you for listening. I’m happy to answer questions.

      Q: What is unique about evaluating in a specific domain like medicine versus general evaluation?

      Moderator: Thank you, Yuexing, for the engaging presentation. We’ll open the floor for questions. I’ll start with one. You’ve described work on AI in a specific domain, medicine, and now more general evaluation. What is unique about working in a specific domain such as medicine compared with general evaluation? What lessons carry across domains?

      Yuexing Hao: During my PhD, I spent a lot of time thinking about this. I started in human-computer interaction, trying to understand user experience in healthcare from physicians, clinician annotators, patients, and underrepresented patient groups. I realized that one problem is the lack of good evaluation and good benchmarks. We need more open-source benchmarks to evaluate the capabilities of different language models.

      Subjective experiences are especially hard to evaluate. Is a system trustworthy, proactive, or polite? People have different standards, and those qualities are difficult to quantify. That made me think about moving from a specific domain such as healthcare to more general domains. Healthcare is sensitive and requires high accuracy. Starting there helps us understand the challenges, and applying those lessons to other domains, especially software testing, is exciting. Healthcare remains a valuable place to study the difficulties of evaluation.

      Q: How do you identify true signal versus stylistic noise in clinicians’ edits?

      Audience member: I’m an emergency physician. USMLE-style questions are relatively easy to evaluate because there is a known answer. Real cases are open-ended. For example, we use vision-language models to draft radiology reports, but it’s difficult to evaluate them. Radiologists make stylistic edits, and terms such as “borderline cardiomegaly” and “mild cardiomegaly” can mean nearly the same thing. Templates also introduce noise. During reinforcement learning, it can be hard to know whether a change reflects a true finding or a stylistic preference. How do you identify the true signal in clinicians’ edits?

      Yuexing Hao: That’s a very good question. Sometimes evaluation is as much art as science, much like clinical diagnosis. It can be hard to say exactly which piece of evidence mattered most until another correlated factor appears.

      MedPAIR is an early benchmark for understanding the differences between clinicians and language models, and we agree that other methods may evaluate this better. We are also working on simulated users, but we are not trying to simulate experts. We mainly simulate members of the general population interacting with tools. Extending simulated users to sensitive areas such as healthcare, law, or finance requires much more consideration.

      Healthcare is not a formula. A diagnosis does not result from assigning one fixed weight to each factor. It is dynamic and depends on time, age, and many other factors. That makes it difficult to use simulated users for these tasks. We may need to start in more general domains, such as e-commerce and advertising, then expand carefully. I appreciate the question.

      Moderator: Thank you. Any other questions about evaluating real-world human experience?

      Q: Do physician-guided signals hurt newer models, and how do you keep up with changing benchmarks?

      Audience member: First, congratulations on the work and on open-sourcing it. I ran some recent models, including GPT-5.6. In my experiment, physician-guided signals and sentences did not help the newer models. They actually reduced performance across several benchmarks. The signals seemed more helpful for weaker models. Is that the pattern you observed?

      I also have a broader question. Benchmarks and evaluation methods can become outdated very quickly. How do you deal with that pace of change? For example, Epoch AI recently published work identifying flawed benchmarks. Who verifies the verifier?

      Yuexing Hao: Thank you. I like both questions. On the first one, evaluation should provide more than a final answer or a single quantitative score. The chain of thought, reasoning process, and order in which a model interprets information also matter. Researchers should ask why a model performs better and what signals it captures, rather than assuming a higher score means it is better.

      For MedPAIR and for model and agent evaluation more broadly, we’re thinking about how evaluations can inform the next round of training and the next model release. Dense reward signals can be useful. MedPAIR involves tedious work: labeling each sentence, aggregating correct labels through majority vote, and releasing the data. We also found noise in the annotations. Some annotators may not specialize in USMLE questions and may answer based on intuition rather than expertise. That limits the reward signal.

      On the broader question, building a static benchmark can be limiting in this era. Models can learn to optimize a score or exploit a benchmark without learning the trajectory that would generalize to the real domain. MatrAIx is trying to make evaluation dynamic. Simulated users with different personas interact with a model in a task environment with defined goals. We evaluate task completion and accuracy, but also user experience and whether the model provides a rationale at each step.

      Dynamic evaluations can also contain noise. Long-tail users may never have used a language model and may prompt it in ways that differ from experienced users. We need to account for that diversity and for the nuances of human behavior.

      Human behavior can change with context. Someone might say they would not buy a Coca-Cola after a two-dollar price increase, but feel differently when thirsty or after an earthquake. We put persona agents in a sandbox with a defined goal and task so we can study behavior in context. We think this can provide a more comprehensive and diverse evaluation.

      Q: Do you need flagship models to run 8.3 billion agents?

      Audience member: I’m interested in the technical side of your startup. Creating 8.3 billion agents is a large undertaking. Do you need flagship models, or can you use lower-cost models?

      Yuexing Hao: We use a lot of GPUs to construct the 8.3 billion virtual personas. We’re also working toward a Guinness World Record by simulating 8.3 billion users answering the same question.

      For a typical software evaluation, though, you do not need to involve 8.3 billion personas. You identify the user segments you care about. If you are developing a social app, for example, you might target people aged 18 to 45, mostly women, living in California. We use those segment definitions with a retrieval system to identify which personas are relevant to the task.

      There are edge cases. In a small country where we have little data, it may be harder to simulate personas. Our business model includes a data flywheel: we work with local populations and use their feedback to improve alignment. If our simulation says 61 percent but reality is 36 percent, that is something we want to correct.

      Q: How do you measure subjective attributes like trustworthiness?

      Audience member: For subjective attributes, it seems impossible to measure everything directly. Are you using proxy metrics?

      Yuexing Hao: Yes. Attributes such as proactivity and trustworthiness are hard to define. Persona agents can act as language-model evaluators and provide a rationale for a rating, such as why they gave a two out of five for trustworthiness. It is still a difficult problem.

      One approach is to ask a language model to judge whether the output meets a criterion. Another is to break a subjective attribute into more objective dimensions. For trustworthiness, for example, you could fact-check the content for hallucinations or factual errors. We are trying to evaluate at that more detailed level.

      Q: How do you gather persona information and match it to a task?

      Moderator: There’s a follow-up in the Zoom chat: How do you gather information about the personas and ensure they match a task?

      Yuexing Hao: We use two strategies. One is to retrieve information from real-world populations. We use sources such as Wikipedia, Amazon reviews, and Stack Overflow’s annual survey to extract relevant persona attributes. We have 2.2 million attributes from real-world personas. For the rest, we use the directed, attribute-based graph I mentioned to model relationships and correlations among factors. We want the attributes to make sense together. We’re also trying to involve more people and may collaborate with Snorkel to bring in human annotators who can provide a more comprehensive understanding of persona attributes.

      Q: How might MedPAIR affect clinicians, trainees, and the medical field?

      Moderator: One final question before we break out for discussion. You’ve talked about MedPAIR’s potential impact on model developers. How might the work affect clinicians, medical trainees, or the medical field itself?

      Yuexing Hao: When we started the project in 2024, we initially thought about perturbations: changing sentence order, changing answer-option order, and adding multimodal inputs. We did not see much change in downstream model performance from some of those experiments, so we pivoted to sentence relevance. We wanted to find out whether relevance was part of the problem and what labels mattered for final predictions.

      We worked with clinician partners to find a way to quantify how clinicians and language models make predictions. At the beginning, we asked annotators to label individual words or phrases they thought were relevant. We found that difficult to quantify. Treating a sentence as the unit of annotation made majority voting easier. I’m happy to share more about the failure cases. I also write about the behind-the-scenes work on my blog, including the blood and sweat that goes into these projects.

      Closing

      Moderator: Thank you, Yuexing, for your work on MedPAIR through Snorkel’s Open Benchmarks Grants program and for your presentation. We welcome everyone to connect with her and with one another. Thank you all.


      Previous Reading Groups

      FAQ

      Q: What is MedPAIR?

      MedPAIR (Medical Dataset Comparing Physicians and AI Relevance Estimation and Question Answering) is a dataset of sentence-level relevance labels on clinical QA questions, collected from physician trainees and board-certified physicians and compared with labels from LLMs. It was presented as an oral at the NeurIPS 2025 Workshop on Socially Responsible and Trustworthy Foundation Models.

      Q: Does a correct answer mean an LLM reasoned like a physician?

      Not necessarily. MedPAIR shows that LLMs often weight different sentences than physicians, and agreement on relevant sentences was roughly 50 to 60% in the talk. The Reading Group description frames the concern as models selecting correct answers through reasoning physicians would dismiss.

      Q: Does removing irrelevant context improve accuracy?

      For many models, yes. Filtering to physician-labeled relevant sentences improved accuracy for physician trainees and most LLMs, but not for every model: Yuexing noted at least one MedGemma result that got worse. Her practical takeaway is that shorter, sharper context can help.

      Q: What does the spurious rate measure?

      It measures how sensitive a model’s answer is to which sentences it sees, meaning how often a question answered correctly with the full context is answered incorrectly once only a selected subset of sentences is kept. In the talk, 46.3% of Qwen 72B’s previously correct answers flipped under physician-relevant-only context, versus 14.9% for Llama 70B.

      Q: Is MedPAIR open source?

      Yes. The labels from physician trainees and LLMs are released at medpair.csail.mit.edu, and the dataset is on Hugging Face.

      Speaker bio

      Yuexing Hao is co-founder and COO of MatrAIx, which builds simulation-based evaluation for AI agents. She was previously a postdoc in MIT’s Healthy ML Group and holds a PhD from Cornell University. Her research focuses on AI for healthcare and human-computer interaction. She is also a researcher at Microsoft and an MIT EECS postdoctoral associate.

      Join us for our next Snorkel AI Reading Group, where a researcher presents a recent paper live and takes audience questions. See upcoming sessions and RSVP at snorkel.ai/reading-group.

      Share this article

      Recommended articles

      View all articles
      Image
      RL environments for LLM agents: Design, rewards, and validation
      TLDR: An agent, by definition, can take actions and is more than just a language model. The model is only one part of the system, and is only one element that you are training. The environment determines what the agent can observe, what it can change, which actions are available, and what behavior receives a reward. For a coding agent,
      September 28, 2026
      •
      Aryan Kargwal
      ,
      Jonathan Schlosser
      Image
      Opus 5.5 vs Opus 5 vs Fable 5.1: Coding Benchmark Results
      We’ve now run three generations of frontier models through the same expert-created terminal-bench style task set: Fable 5.1 in September, Opus 5 in July, and, now Opus 5.5. For this analysis, the task set contained 200 trajectories, of 24 tasks, with every failure traced to a judge-confirmed root cause. A generational climb On the task set, pass@1 is 61.5% for
      September 23, 2026
      •
      Ankit Aich
      ,
      Jonathan Schlosser
      series-e-blog
      Data 2.0 and the research era of AI data
      Today, I’m excited to announce Snorkel’s $350M Series E financing at a $3.5B valuation, led by Insight and S32, with participation from Third Point, March, Blumberg, Allegis, Standard VC, Frontline, and existing investors Addition, Lightspeed, Greylock, GV, P7, Wells Fargo, Walden Catalyst Ventures, and Factory.
      September 22, 2026
      •
      Alex Ratner
      Image

      Join our newsletter

      For expert advice, the latest research, and exclusive events.
      By submitting this form, I acknowledge I will receive email updates from Snorkel AI, and I agree to the Terms of Use and acknowledge that my information will be used in accordance with the Privacy Policy.