
NVIDIA unveils NV-Reason-CT, an open 3D CT model with radiologist-style reasoning
NVIDIA introduced NV-Reason-CT, a vision-language model built for full 3D CT analysis that generates structured reports and radiologist-style chain-of-thought reasoning, with open checkpoints on Hugging Face.
NVIDIA has introduced NV-Reason-CT, a vision-language model (VLM) purpose-built for analyzing three-dimensional computed tomography (CT) scans. It produces structured diagnostic reports and step-by-step reasoning that mirrors how radiologists work through a study, the company said in a blog post on September 23, 2026.
The release extends to volumetric imaging the reasoning approach NVIDIA pioneered with NV-Reason-CXR, a chest X-ray model whose multireader study was accepted at RSNA 2026 and confirmed radiologist time savings while maintaining diagnostic accuracy.
Why 3D CT needs its own approach
A single abdominal CT study can contain 300–600 axial slices, encoding anatomy across three spatial dimensions. Standard VLMs treat image input as a 2D grid of tokens, so processing a volume as a stack of independent frames discards the relationships between slices that make masses, effusions and infiltrates clinically meaningful.
NVIDIA points to two further gaps: models that spot an abnormality often output a label without explaining why, and most systems lack the multiturn dialogue radiologists need to probe a finding.
A 3D transformer paired with a 4B language model
The architecture combines a full 3D vision transformer encoder with a Qwen3.5-4B language model. CT volumes are resampled to 192³ voxels at 2 mm isotropic resolution and split into non-overlapping 8×8×8 patch tokens — 13,824 vision tokens in all, passed to the language model with their 3D grid coordinates. A 3D MRoPE scheme keeps spatial relationships intact through the LLM layers. All weights were retrained end-to-end.
Training ran in two stages: supervised fine-tuning on roughly 550,000 structured QA examples for chest and abdomen, including radiologist-dictated reasoning annotations, then reinforcement learning with GRPO and an anatomy-aware reward.
Benchmarks and clinical review
On CT-RATE, the leading public benchmark for 3D CT understanding, NV-Reason-CT reports a Macro-F1 of 0.614 and a Macro-AUROC of 0.871, ahead of 3D contrastive models such as VoxelFM (0.581/0.870) and Pillar-0 (0.544/0.861), as well as ClinFusion-8B (0.442) and MedGemma 1.5 (0.303). NVIDIA says it is the first time a single open model has been competitive at both CT classification and report generation.
NIH radiologists reviewed the outputs; Baris Turkbey, M.D., F.S.A.R., a senior clinician at the National Institutes of Health, said the model's step-by-step reasoning reflects how radiologists actually think through a CT study, and that reviewing its thought process rather than only its conclusions makes the findings possible to trust and act on.
What “open” covers
NV-Reason-CT is an open research and development foundation, not an autonomous diagnostic system or a cleared clinical product. Checkpoints are available on Hugging Face, and the GitHub repository includes inference scripts, training configurations and post-training recipes. It joins the NVIDIA Medical AI family alongside NV-Generate-CTMR for synthetic volumes and NV-Segment-CTMR for segmentation.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.