*** KONVENS 2026 Call for participation *** https://konvens2026.uni-hamburg.de/
We would like to invite you to participate on site at the Konferenz zur Verarbeitung natürlicher Sprache (KONVENS) 2026, organized under the auspices of the GSCL, the DGfS-CL, the ÖGAI, and SwissNLP. This year’s KONVENS will take place in Hamburg, September 14 – 17 under the special theme “Context Matters: NLP Beyond Text”. The conference will include a diverse program including talks by our two keynote speakers:
– KEYNOTE 1 –
Tuesday, 15 September 2026 · 13:30 –14:30 · Lecture Hall D
The Myth of Ground Truth: Why Context and Disagreement Matter for NLP Barbara Plank
Abstract: Natural language processing has made remarkable progress. Yet much of this progress still rests on a powerful simplifying assumption: that for every input there exists a single correct answer, a single ground truth, to be recovered independent of who is asking, who is answering, and in what context.
In this talk, I argue that this assumption is often at odds with how language actually works, and with the complex, ambiguous, and inherently human environments in which language technology operates. Meaning is filled in by context: by what a sentence implies rather than states explicitly, by who interprets it and what perspective, language, dialect, or background they bring. Drawing on research on human label variation, ambiguity, and the evaluation of foundation models, I show that this variation is natural, not noise to be eliminated, but signal to be understood. I close by considering what this means for generative AI and alignment.
Human-centered NLP is multi-faceted: building inclusive systems around people means building them around their disagreement, not despite it.
– KEYNOTE 2 –
Wednesday, 16 September 2026 · 09:00 –10:00 · Lecture Hall D
Evaluation in Context: Humans, Machines, and the Crisis of Validity Valentin Hofmann
Abstract: Language model evaluation is facing a crisis of validity: models increasingly excel on established benchmarks, yet benchmark performance can be a poor indicator of how systems behave in the wild. In this talk, I will argue that making evaluation meaningful requires taking context seriously — not only the context of the humans who use and are affected by language models, but also the context of the models being evaluated.
I will first examine social bias evaluation as an example of core validity issues in current practices and discuss efforts to build benchmarks that more closely reflect the ways in which people actually use language models as well as the harms they experience. Turning from humans to machines, I will then show how the value of an evaluation item depends on the model taking the test, and how this insight can be leveraged through psychometric methods to improve validity. I will conclude by making the case for a future of language model evaluation that takes into account both the humans around our systems and the machines in front of us.
Registration: https://konvens2026.uni-hamburg.de/index.php/registration/