LLM Evaluation NLP · Pragmatics Computational Linguistics

Roland
Mühlenbernd

// ML Researcher · Architect of Language Models

I build ML models of language — and increasingly, I evaluate whether today's LLMs actually understand it. My current work develops novel calibration metrics (ESR, CDS) to benchmark frontier models (GPT, Claude, Gemini) on fine-grained social and pragmatic tasks.

My background spans Computer Science, Media Studies, and Computational Linguistics (BSc–MSc–PhD). I bring 15+ years of rigorous modeling — game theory, multi-agent RL, probabilistic NLP — to the question that matters most in applied AI: does the model actually get what humans mean?

I'm actively seeking roles in NLP research, LLM evaluation, or applied language AI where linguistic depth meets engineering ambition.

🏅 ERC Seal of Excellence 💶 €75K Research Grant (NAWA) 🎤 Invited Speaker, Stanford
Roland Mühlenbernd
40+
Publications
50+
Presentations
15+
Years Research
25+
Courses Taught
researcher_profile.py
# Roland Mühlenbernd focus = [   "LLM evaluation & calibration",   "pragmatics & social meaning",   "language dynamics", ] stack = {   "ml": ["PyTorch", "HuggingFace"],   "stats": ["Python", "R"], } publications = 40 # and counting
Skills & Domain

Tech and Theory

Technical Stack
  • PyTorch · TensorFlow · scikit-learn
  • HuggingFace Transformers · fine-tuning
  • LLM evaluation · prompt engineering
  • Python (expert) · R · Julia · C++
  • spaCy · NLTK · Pandas · NumPy
  • Git · JupyterLab · Google Colab
  • Statistical modeling · Bayesian inference
Linguistic Expertise
  • Pragmatics · politeness theory · register
  • Language evolution & historical change
  • Sociolinguistics · network variation
  • Formal semantics · game-theoretic pragmatics
  • Experimental linguistics · behavioral data
  • Morphology · grammaticalization
  • Computational sociolinguistics
Current Work

What I'm Working On

Active research at the intersection of NLP, LLM evaluation, and computational pragmatics — plus the datasets, code, and demos it produces.

● Active

LLM Calibration Metrics for Social Meaning Tasks

Developing ESR and CDS metrics to evaluate GPT, Claude, and Gemini on politeness, precision, and register · published in CMCL 2026 proceedings · paper · code & data

✓ Published
● Active

Motive-Mediated Social Inference in LLMs

Follow-up to CMCL 2026 · tests 10 models · cross-model calibration differences trace to motive-to-judgment mapping quality, not to which motives surface in attribution · submitted to EMNLP 2026, currently under review

Under review · EMNLP
● Active

PRISMA — a Chatbot That Profiles You Back

A Gradio demo (GPT-OSS 120B via Groq) that rates you on six social-perception dimensions turn by turn while it chats with you · built on the CMCL 2026 / EMNLP (under review) research · try it live · code

▶ Live Demo
● Active

imprecision-bench: A Multimodal Pragmatics Benchmark

475 human time-reference productions across 12 clock states × 2 pragmatic contexts, with a peer-reviewed RSA baseline (r² ≈ 0.97) · tests whether LLMs and VLMs calibrate precision to social context · dataset on HF · code

◆ Public Dataset
● Active

Probabilistic Pragmatics: Speaker Models & Strategic Requests

Extending RSA and game-theoretic frameworks to model strategic linguistic choices — (im)precision, politeness, register, and direct vs. indirect requests — as sources of social meaning beyond literal content · part of SFB 1412 Register project at ZAS Berlin

In progress
Recent writing →
June 25, 2026
When Your Chatbot Gets Judgmental…
Meet PRISMA — the LLM that elucidates social meaning.