// ML Researcher · Architect of Language Models
I build ML models of language — and increasingly, I evaluate whether today's LLMs actually understand it. My current work develops novel calibration metrics (ESR, CDS) to benchmark frontier models (GPT, Claude, Gemini) on fine-grained social and pragmatic tasks.
My background spans Computer Science, Media Studies, and Computational Linguistics (BSc–MSc–PhD). I bring 15+ years of rigorous modeling — game theory, multi-agent RL, probabilistic NLP — to the question that matters most in applied AI: does the model actually get what humans mean?
I'm actively seeking roles in NLP research, LLM evaluation, or applied language AI where linguistic depth meets engineering ambition.
Active research at the intersection of NLP, LLM evaluation, and computational pragmatics — plus the datasets, code, and demos it produces.
Developing ESR and CDS metrics to evaluate GPT, Claude, and Gemini on politeness, precision, and register · published in CMCL 2026 proceedings · paper · code & data
Follow-up to CMCL 2026 · tests 10 models · cross-model calibration differences trace to motive-to-judgment mapping quality, not to which motives surface in attribution · submitted to EMNLP 2026, currently under review
A Gradio demo (GPT-OSS 120B via Groq) that rates you on six social-perception dimensions turn by turn while it chats with you · built on the CMCL 2026 / EMNLP (under review) research · try it live · code
475 human time-reference productions across 12 clock states × 2 pragmatic contexts, with a peer-reviewed RSA baseline (r² ≈ 0.97) · tests whether LLMs and VLMs calibrate precision to social context · dataset on HF · code
Extending RSA and game-theoretic frameworks to model strategic linguistic choices — (im)precision, politeness, register, and direct vs. indirect requests — as sources of social meaning beyond literal content · part of SFB 1412 Register project at ZAS Berlin