I'm a Senior Quality Engineer based in Hamburg, Germany. 10+ years across QA, test automation, and software engineering. Currently working on LLM and AI systems evaluation.
Improving quality is risk reduction: building for testability, getting feedback early, and asking whether what we measure reflects what matters.
-
review-sentiment-eval An evaluation suite for an LLM-based review analysis system. Covers summarization faithfulness and prompt injection robustness, with calibrated LLM-as-judge scoring and an adaptive baseline where the check only fails when the injection changes the output.
-
weathershopper-playwright-test-suite E2E coverage of a weather-driven shopping flow. Notable for dynamic product selection over hardcoded names, and handling the site's intentional 5% payment failure with a single retry, taking false alarms from 5% to 0.25%.
Doctoral research in audio signal processing, separating overlapping sources from a single recording. It is an underdetermined problem: many valid outputs, no single correct one.


