Nariman
Naderi

Physician Programmer · Gastroenterology AI Research Python · LLM evaluation · Computer vision

I trained as a physician.
Now I test whether AI can be trusted at the bedside.

I write the code behind my research.
Then I find out where the models break.

Read as
Nariman Naderi, physician programmer and medical AI researcher

About

I study where medical AI falls short

I measure where models fail on clinical data

I am a physician researcher at the Gastroenterology and Liver Research Center, Taleghani Hospital.

My questions come from clinical work. Can a language model tell a gastroenterologist when it is unsure? Can it reconcile IBD guidelines that disagree? Do the CT datasets we train on look like the patients we actually see?

I write the code behind my own research, mostly in Python.

My work covers LLM evaluation and confidence elicitation, retrieval-augmented generation over clinical guidelines, zero-shot vision-language models benchmarked against trained ML models, and multimodal CNNs on EEG spectrograms.

I have always been drawn to problems that refuse a quick answer, especially when solving them can lead to something useful.

When a model fails, I want to know which assumption broke first.

How I got here

From code, to medicine, to both

git log --reverse

  1. Before medicine

    I learned to code first.

    Programming came before medical school. It taught me to break a hard problem into steps small enough to check.

    Programming came first, before medical school. Small programs, one bug at a time.

  2. Shahid Beheshti

    Then medicine took over.

    At Shahid Beheshti University of Medical Sciences, clinical training gave me the questions I still work on.

    Medical school at Shahid Beheshti. The code went quiet, but the habit of asking how a system fails stayed.

  3. 2022

    The two came back together.

    Midway through my MD, I started writing the code for my own research. That year, our COVID-19 cohort study from Tehran was published.

    Picked Python back up seriously. I have written the code for my own AI-in-medicine research ever since.

  4. 2025

    MD, then gastroenterology AI.

    I finished my MD. Today I research at the Gastroenterology and Liver Research Center, asking whether AI can be trusted with GI patients.

    MD done. Now at the Gastroenterology and Liver Research Center: LLM confidence, RAG over guidelines, vision-language models on colonoscopy.

  5. Now

    Two questions on my desk.

    Both start with the same worry: would this still work for my patients?

    Both are evaluation problems before they are modeling problems.

Publications

Selected papers

Across generations, sizes, and types, large language models poorly report self-confidence in gastroenterology clinical reasoning tasks.

npj Gut and Liver · First author (equal contribution)

Can a language model tell you when it is unsure about a GI case? Mostly, no.Confidence reporting compared across LLM generations, sizes, and types.

Vision language models versus machine learning models performance on polyp detection and classification in colonoscopy images.

Scientific Reports

Can a general-purpose AI model find and classify colon polyps?Zero-shot vision-language models benchmarked against trained ML models on colonoscopy images.

State of abdominal CT datasets: A critical review of bias, clinical relevance, and real-world applicability.

PLOS Digital Health

Do public abdominal CT datasets look like the patients we treat?Critical review of dataset bias, clinical relevance, and real-world applicability.

Using large language models to integrate international IBD guidelines: A retrieval-augmented generation approach.

Colorectal Disease

Can AI merge international IBD guidelines into one consistent answer?Retrieval-augmented generation over international guidelines.

Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs.

EXTRAAMAS · Lecture Notes in Computer Science · First author

Does the way you ask a language model change how right, and how sure, it is?Prompting techniques compared for accuracy and confidence elicitation.

The challenge of uncertainty quantification of large language models in medicine.

arXiv

When should a clinician trust a language model's answer?Perspective on uncertainty quantification methods for medical LLMs.

Alzheimer's Disease Stage Classification via Multimodal CNN on EEG Spectrograms and Cube-Drawing Images.

medRxiv

Can an EEG and a cube-drawing test stage Alzheimer's disease?Multimodal CNN on EEG spectrograms and drawing images.

Independent and Combined Association of Overweight and Obesity and Metabolic Dysfunction–Associated Steatotic Liver Disease (MASLD) with 10-Year Atherosclerotic Cardiovascular Disease Risk in Patients with Type 2 Diabetes.

Iranian Journal of Nutrition Sciences and Food Technology · First author

How do excess weight and fatty liver add up to heart risk in type 2 diabetes?Independent and joint association analysis against 10-year ASCVD risk.

Association of Lifelines Diet Score (LLDS) with Risk of Pancreatic Steatosis: A Case-Control Study.

Iranian Journal of Endocrinology and Metabolism · First author

Is diet quality linked to fat in the pancreas?Case-control study of a diet score against pancreatic steatosis.

Assessing the effect of remdesivir alone and in combination with corticosteroids on time to death in COVID-19: A propensity score-matched analysis.

Journal of Clinical Virology Plus

Did remdesivir, alone or with steroids, change time to death in COVID-19?Propensity score–matched time-to-event analysis.

Epidemiology of COVID-19 in Tehran, Iran: a cohort study of clinical profile, risk factors, and outcomes.

BioMed Research International

Who became sickest with COVID-19 in Tehran, and what predicted it?Cohort study modelling risk factors and outcomes.

Contact

Have a clinical question AI could help answer?

Have a model that needs a clinician's eye?

I welcome collaborations on medical LLM evaluation, clinical guideline synthesis, and medical imaging. If you have a research question or study idea, email me. If you are building or evaluating a model on medical data and want someone who can read both the code and the chart, email me.