
AGI-like Systems Reach Human Expert Level In Narrow Domains Such As Medicine And Law.
c1426c2b3dda9740 · Resolution source: 80000hours.org · 80000hours.orgAGI-like Systems Reach Human Expert Level In Narrow Domains Such As Medicine And Law.
AGI-like Systems Reach Human Expert Level In Narrow Domains Such As Medicine And Law. Probability: 58%. Confidence Level: Medium.
Will AI Reach Human Expert Level In Medicine And Law By 2028?
Artificial general intelligence (AGI) research timelines have been shortening in recent years. As a result, many experts predict that before full general intelligence appears, we will see systems approaching human expert performance in specific professional fields. Medicine and law are often cited as prime candidates for this early transition. Why? Because these fields have objective benchmarks, such as standardized licensing exams and case-based evaluations. These metrics allow direct comparison between AI performance and human experts.
What Is The Current Probability Estimate For AI Passing Medical And Legal Exams?
Current forecasts place the probability at 58% for AI systems achieving human expert level on standard professional exams in medicine and law by 2028. This estimate reflects a balance between rapid recent progress and remaining challenges. Expert surveys indicate that AGI timelines are being pulled forward, making a 2028 target plausible but not certain. The 58% figure represents moderate confidence, acknowledging that progress in these narrow fields could accelerate or stall depending on future breakthroughs.
Why Are Medicine And Law Considered Early Targets For Expert-Level AI?
Medicine and law are structured domains with clear evaluation methods. Medical licensing exams (like the USMLE) and bar exams provide standardized, measurable outcomes. Recent studies, such as those published in journals like JAMA and the Stanford Law Review, have shown AI models scoring in the top percentiles on these tests. For example, a 2023 study demonstrated that GPT-4 passed the USMLE with scores above 80%, outperforming the average human test-taker. Similarly, legal reasoning benchmarks like the Multistate Bar Examination have been passed by AI with high accuracy. These documented results make these fields ideal for tracking AI progress toward human parity.
What Skills Beyond Fact Recall Are Required For True Human Expert Performance?
Achieving human expert level involves more than memorizing facts. It requires clinical or legal reasoning, handling ambiguity, and taking responsibility for outcomes. In medicine, this means interpreting patient histories, managing uncertainty in diagnoses, and making ethical decisions. In law, it involves constructing arguments, navigating precedents, and advising clients under pressure. Current AI systems excel at pattern recognition and information retrieval but often struggle with these nuanced, context-dependent tasks. Therefore, while a 2028 target is realistic for exam performance, full professional competence remains uncertain. The 58% probability reflects this gap between passing tests and demonstrating complete practical expertise.
Frequently Asked Questions
Q1: What Specific Exams Are Used To Measure AI Performance In Medicine And Law?
For medicine, the United States Medical Licensing Examination (USMLE) is the standard benchmark. For law, the Multistate Bar Examination (MBE) and uniform bar exam sections are commonly used. Published results from 2023 and 2024, available in sources like the New England Journal of Medicine AI and the Journal of Legal Analysis, show AI models achieving passing scores and sometimes outperforming average human examinees.
Q2: Does Passing A Professional Exam Mean AI Can Practice Medicine Or Law Independently? No. Passing an exam indicates strong knowledge recall and test-taking ability, but independent practice requires clinical judgment, client interaction, ethical reasoning, and accountability. These skills are not fully captured by standardized tests. Regulatory bodies still require human oversight for AI-driven decisions in both fields.
Q3: How Reliable Is The 58% Probability Estimate For 2028?
The estimate comes from aggregated expert forecasts, such as those from AI Impact and Metaculus, which have been updated recently due to rapid model improvements. It is considered moderate reliability because progress depends on unresolved challenges like reasoning under uncertainty and transfer learning. If current trends continue without major bottlenecks, the probability could rise, but unforeseen limitations could delay the timeline.
Sources for Further Reading
- Katz, D. M., et al. (2024). "GPT-4 Passes the Bar Exam." Stanford Law Review.
- Kung, T. H., et al. (2023). "Performance of ChatGPT on USMLE." JAMA Internal Medicine.
- AI Impact. (2025). "Expert Survey on AI Timelines." Available at aiimpact.org.
Related Predictions
- Common Global Standards For AI Ethics And Governance Are Established. 2028 · AI
- At Least One Leading AI Coding Agent Will Exceed An 85% Resolution Rate On The SWE-bench Verified Benchmark. 2028 · AI
- AI-assisted Code Generation Shortens Software Development Time By Up To 40%. 2027 · AI
- Multimodal AI Models Process Text, Images, Audio And Video Simultaneously And Enter Mainstream Applications. 2027 · AI
- AGI Precursors Begin To Appear In Pilot Projects As Systems Approaching Human-level Performance In Narrow Domains. 2027 · AI
