π₯ ClinFusion-32B
A Vision-Centric Multimodal LLM for Holistic Medical Understanding (32B Parameter Flagship)
ClinFusion-32B is the flagship, high-performance model of the ClinFusion multimodal large language model series. Built on Qwen3-VL-32B-Instruct, ClinFusion-32B is engineered for complex clinical reasoning, high-fidelity medical report generation, and fine-grained visual-textual instruction-following.
By fusing DINOv2 and ConvNeXt vision encoders with a 32-billion parameter language backbone through our custom Cascade Spatial-Aware Locality Fusion operator, ClinFusion-32B delivers unmatched comprehension of medical cases, approaching and sometimes exceeding proprietary models like GPT-5.2 and Gemini-3-Flash.
π Key Features of ClinFusion-32B
- Compositional Vision Architecture: Combines DINOv2 (for dense spatial semantics) and CLIP-ConvNeXt (for structural and holistic features) with Qwen3-VL, optimized via a Cascade Spatial-Aware Locality Fusion operator.
- Native 2D & 3D Multimodal Support: Seamlessly processes standard 2D medical images (e.g., X-ray, Ultrasound, Histology) and native 3D medical volumes (NIfTI
.nii.gzCT/MRI files). - Flagship Clinical Reasoning (32B Backbone): Highly capable of synthesizing information across long reports, complex textual medical benchmarks, and intricate multi-image patient cases with superior clinical accuracy.
- SOTA Benchmarks & Agentic Workflows: Sets a new state-of-the-art across 2D/3D benchmarks, outperforming proprietary models (e.g., GPT-5.2) on 13 out of 16 suites, while natively supporting tool-use for clinical workflows.
π Quick Start & Usage
To set up the environment, run inference, or evaluate ClinFusion-32B on your own data, please refer directly to our official GitHub Repository.
The GitHub repository provides comprehensive, step-by-step instructions for:
- One-click installation using
uv. - Data preparation (supporting 2D images, text-only, and 3D NIfTI volumes).
- Multi-GPU / Multi-Node deployment for executing the 32B flagship model.
- Model inference and evaluation scripts.
π Evaluation and Performance
ClinFusion-32B sets a new state-of-the-art across a comprehensive suite of 2D and 3D multimodal medical benchmarks.
- Outperforms leading open-source medical models (e.g., Hulu-Med, Lingshu) on 20 out of 24 benchmarks.
- Outperforms proprietary foundation models (GPT-5.2, Gemini-3-Flash) on 13 out of 16 core evaluation suites.
π Citation
If you find ClinFusion useful in your research, please consider citing:
@article{yuan2026ClinFusion,
title={ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding},
author={Yuan, Hangjie and Qian, Yichen and Tang, Zhiwei and Xu, Xianzhe and Wu, Lirong and Yang, Sicheng and Wang, Jinwang and Wang, Pengju and Zeng, Zhitao and Han, Yizeng and Xing, Yan and Luo, Shengxuan and Feng, Tao and Xie, Qing and Yao, Weigen and Yang, Yi and Liu, Zuozhu and Tang, Jiasheng and Wang, Shaocheng and Wang, Jitao and Dong, Jiahong and Chen, Weihua and Xu, Feng and Wang, Fan},
journal={arXiv preprint arXiv:2607.24743},
year={2026}
}
π Contact & Acknowledgements
Built with β€οΈ by Alibaba DAMO Academy. Special thanks to the open-source community behind Qwen3-VL, DINOv2, and OpenCLIP.
- Downloads last month
- 218