About
I recently completed my B.Comp. (Honours) in Computer Science with a specialisation in Data Science at the College of Computing and Data Science, Nanyang Technological University, Singapore.
Since 2024 I have been a research collaborator at the State Key Laboratory of Respiratory Disease, Guangzhou Medical University, working with Prof. Jianxing He's thoracic surgery and thoracic oncology group on visual foundation models for lung-cancer diagnosis from histopathology and computed tomography.
My research asks a single question from several angles: how can general-purpose foundation models be made to reason reliably over medical data they were never fine-tuned on? Concretely, I work on (i) training-free multimodal frameworks that inject CT and whole-slide-image evidence into long-context LLMs as aligned pseudo-tokens, (ii) process reward modelling that supervises each step of a biomedical reasoning chain rather than only its final answer, and (iii) source-free domain adaptation that recovers segmentation accuracy across staining protocols without source data or new annotation.
Along the way I have been a research intern at POSTECH (DPNM Lab, with Prof. James Won-Ki Hong), a research assistant at NTU's S-Lab for Advanced Intelligence and School of Social Sciences, and an independent open-source developer of speech-synthesis tooling used by content creators.
I am actively looking for PhD opportunities in multimodal and medical AI. If you are recruiting students, or would be interested in collaborating on any of the above, I would be glad to hear from you by email.
Selected Publications
View All →Process Reward Model as Cross-Task Learner: An Extension of Vision-Language Model to Zero-Shot Medical Reasoning
Jiahua Zhang, Yidong Tian, Jinghao Liang, Dianhan Lin, Yiwen Cai, Yuanqing Liu, Zishan Huang, Jingchun Ni, Jianxing He†
Proceedings of the 34th ACM International Conference on Multimedia (ACM MM)
MedTRACE trains a lightweight process reward model once, then reuses it to steer a frozen vision-language model across unseen CT and whole-slide-image tasks with no gradient updates at deployment.
Can LLM Understand Medical Imaging with Long Context?
Jiahua Zhang, Jinghao Liang, Yiwen Cai, Jingchun Ni, Jianxing He
Proceedings of the IEEE International Conference on Multimedia and Expo (ICME)
Oral at ICME 2026. Frozen CT/WSI encoders are compressed and optimal-transport-aligned into a long-context LLM as pseudo-tokens, giving few-shot diagnosis, staging and response assessment with no fine-tuning on either side.
News
🎉 Process Reward Model as Cross-Task Learner was accepted at ACM Multimedia 2026.
🎓 Graduated from NTU with a B.Comp. (Honours) in Computer Science, specialising in Data Science.
🎉 Can LLM Understand Medical Imaging with Long Context? was accepted at IEEE ICME 2026 as an oral.
📄 Wrapped up the BioMedPRM internship on process reward modelling for biomedical reasoning (NTU CAO Exchange Programme, Shenzhen).
🎤 Presented DiffiT-HSFDA as an oral at ICIC 2025, published in Springer LNCS.
📍 Completed the POSTECH Summer Program at the DPNM Laboratory with Prof. James Won-Ki Hong.