Empowering Indian Healthcare AI with High-Quality Data
At Eka Care, we believe that advancing healthcare in India requires AI solutions built specifically for Indian populations. To support this vision, we share anonymized, expert-annotated healthcare datasets with the research community. All datasets are in the EkaCare Medical Public Datasets collection on Hugging Face and on the India AI AIKosh platform.Speech
Eka Medical ASR Evaluation Dataset
- A curated collection of 3,900+ English and Hindi medical speech recordings designed to benchmark and improve speech-to-text performance for Indian healthcare use cases.
- Built with diverse accents and real clinical terminology, with per-term medical entity annotations and a published leaderboard on the dataset card. Used by Google Health, ElevenLabs and academic groups for fine-tuning and benchmarking.
- AIKosh | Hugging Face
Clinical Documentation
Clinical Note Generation Dataset
- A multilingual dataset of 156+ transcribed doctor–patient conversations (English, Hindi, Marathi) designed to evaluate AI systems that convert medical dialogues into structured, entity-level clinical records.
- Expert-annotated ground truth JSON and rubric-based LLM evaluation ensure accurate benchmarking of structured note generation for EHR-ready medical documentation.
- AIKosh | Hugging Face
EkaCare Medical History Summarisation
- A curated set of 58 real-world medical cases designed to evaluate AI systems in generating concise, clinically relevant summaries of patient histories.
- Expert-defined rubrics and reference summaries ensure objective assessment of key developments, historical context, and critical care insights from the most recent six months of medical data.
- AIKosh | Hugging Face
Medical Documents
Medical Records Parsing Validation Set
- A curated dataset of 288 de-identified lab reports and prescriptions designed to evaluate AI systems that extract structured data from unstructured medical documents.
- Expert-annotated with rubric-based LLM evaluation, it ensures clinically accurate benchmarking across diverse Indian healthcare document formats.
- AIKosh | Hugging Face
Embeddings & Retrieval
Eka-IndicMTEB
- A multilingual medical embedding benchmark with 2,532 doctor-verified queries across 8 Indic languages, aligned to SNOMED CT for concept-level evaluation.
- Designed to benchmark and improve cross-lingual medical retrieval and semantic search systems across India’s diverse linguistic landscape.
- AIKosh | Hugging Face | Launch blog
Medical Knowledge (MCQ)
Indian Drug MCQA
- 1,512 multiple-choice questions on Indian branded medications across 20+ therapeutic classes, testing whether a model can identify the generic composition from a brand name.
- Includes hard-difficulty variants with alternative distractor strategies and phrasings.
- AIKosh | Hugging Face
Eka NFI MCQA
- 925 multiple-choice questions derived from the National Formulary of India (NFI) 2011, covering indications, contraindications, dosing, precautions, pregnancy safety and drug schedules.
- Evaluates drug pharmacology knowledge as defined by the Indian government reference rather than Western formularies.
- AIKosh | Hugging Face
Clinical Reasoning
Medical Calculator Evaluation Dataset
- 1,066 questions across 358 medical calculators (BMI, GFR, APACHE II, drug dosing and more) in 24 domains, testing whether LLMs can compute numeric medical values from parametric memory and inline arithmetic alone.
- Structured JSON Q&A with numeric comparison and per-row tolerance for evaluation.
- AIKosh | Hugging Face
Indian Protocols-Based Clinical Q&A
- A rubric-graded evaluation set built from Indian and international clinical guideline documents. Each sample is a realistic doctor-side query against a known protocol.
- Rubrics grade both whether the system retrieved the correct guideline content and whether the final answer is clinically complete and safe, stress-testing protocol-grounded question answering.
- AIKosh | Hugging Face
Clinical Knowledge Graphs (BODHI)
BODHI (Bharat Ontology for Disease & Healthcare Informatics) is a set of open, clinician-validated knowledge graphs for grounding healthcare AI in verified medical facts. Available in Neo4j, CSV, PyTorch Geometric and RDF formats. Read the launch blog.BODHI-S — Condition–Symptom Graph
- Maps 779 clinical conditions to 4,037 symptom variants, 39 specialities and inter-condition risk relationships (13,204 relationships in total), with triage levels and demographic likelihood scores on every node.
- Built and validated by expert clinicians and has powered production symptom checking and differential diagnosis across millions of patient interactions in India.
- AIKosh | Hugging Face | GitHub
BODHI-M — Concept–Drug–Lab Graph
- Maps 2,471 SNOMED-coded clinical concepts to 1,186 generic drugs and 812 LOINC-coded lab investigations, organised in a three-level System → Group → Granular hierarchy.
- Supports soft inference at broader levels when a drug or lab result cannot pinpoint a specific disease, and reverse inference of likely conditions from a medication list.
- AIKosh | Hugging Face | GitHub
Population Health
NidaanKosh
- A comprehensive laboratory investigation dataset of 100,000 Indian subjects with 6.8 million+ readings.
- Covers common biomarkers and laboratory values specific to Indian populations, enabling population-level reference ranges and screening research.
- AIKosh | Hugging Face
Spandan
- A large photoplethysmography (PPG) signal dataset of 1 million+ Indian subjects, with raw signals captured from diverse demographic groups across India.
- Essential for developing accurate cardiovascular monitoring algorithms for Indian populations.
- AIKosh | Hugging Face
Why These Datasets Matter
Bridging the Data Gap
Majority healthcare AI models developed today are trained on Western datasets, which may not accurately represent Indian patients’ unique characteristics. These datasets address this critical gap.Enabling Homegrown Innovation
With access to high-quality Indian healthcare data, researchers and developers can build AI solutions tailored specifically to Indian healthcare challenges.Advancing Healthcare Equity
By democratizing access to these datasets, we aim to support broader participation in healthcare AI development across India.How to Use These Datasets
- Access: All datasets are on Hugging Face and the India AI AIKosh platform. A few require accepting the terms on Hugging Face before download.
- Evaluate: Use KARMA – OpenMedEvalKit, our open-source medical AI evaluation framework, to benchmark models on these datasets.
- Documentation: Each dataset card describes the data structure, collection methodology, and suggested applications.
Ethics & Privacy
All datasets have been rigorously anonymized following industry best practices and ethical guidelines. No personally identifiable information is included in any dataset.Research Collaboration
We welcome collaboration opportunities with academic institutions, research organizations, and industry partners. If you’re using our datasets for your research:- Please cite our datasets in your publications
- Share your findings with us — see who is already building on Eka open assets
- Consider opportunities for joint research initiatives
These datasets are provided for research and development purposes. For terms of use and licensing information, please refer to the license on each dataset card.

