I develop reliable language models and NLP methods for medicine, spanning biomedical pre-training, clinical adaptation, and evaluation.
Reliable clinical prediction
Much of what clinicians know about a patient is written in notes rather than coded fields. I develop language-model methods for outcomes such as post-discharge mortality and psychiatric readmission, focusing on confidence estimation, explanation generation, and reasoning over long clinical documents.
Language models for biomedicine
I co-led BioBERT, one of the first language models pre-trained on biomedical literature. Later work focused on making such models usable beyond the lab through compact entity-recognition models and the open-source KAZU framework, developed with AstraZeneca.
Clinical information extraction
I study temporal information, entities, and relations in biomedical and clinical text. This work includes community benchmarks in question answering and information extraction, as well as co-organizing the 2024 and 2025 ChemoTimelines shared tasks on reconstructing cancer-treatment timelines from clinical narratives.


