
I’m a postdoc at TakeLab, University of Zagreb, working on faithful explainability, safety, and controllability of language models.
Previously, I was a postdoc with Yonatan Belinkov at the Technion, on unlearning and faithful explainability of language models, and before that with Iryna Gurevych at the UKP Lab, TU Darmstadt, on the InterText initiative. I did my PhD at the University of Zagreb with Jan Šnajder, and before that worked at the European Commission’s Joint Research Centre on NLP for the Sendai Framework for Disaster Risk Reduction.
I am on the job market for academic opportunities. Check my CV and reach out if you believe me a good fit.
News
- Jul 2026 — "Reasoning Models Know What's Important" accepted to COLM 2026 [paper]
- Jun 2026 — "Evaluating Pluralism in LLMs" accepted to the Pluralistic Alignment workshop @ ICML 2026 [paper]
- May 2026 — "Old Habits Die Hard" accepted to ICML 2026 [paper]
- Apr 2026 — CRISP accepted to ACL 2026 [paper]
- Mar 2026 — I gave talks on "From Internals to Integrity: How Insights into Transformer LMs Improve Safety and Faithfulness" at Bocconi University, IMS @ University of Stuttgart, RTG Neuroexplicit models @ University of Saarland, Data and Web Science Group @ University of Mannheim and Lamarr Institute @ University of Bonn!
- Jan 2026 — ManagerBench accepted to ICLR 2026 [paper & data]
- Jan 2026 — Sequence-repetition paper accepted to Findings of EACL 2026 [paper]
- Nov 2025 — FUR wins the Outstanding Paper Award at EMNLP 2025 [paper]
Older news
- Nov 2025 — PragWorld accepted to AAAI 2026 as an oral [paper]
- Sep 2025 — New preprint: Context Parametrization with Compositional Adapters [paper]
- Aug 2025 — New preprint: CRISP, erasing harmful concepts from LMs via SAEs [paper]
- Jul 2025 — Model-editing intrinsic-features paper accepted to the Interplay workshop @ COLM 2025 [paper]
- Jun 2025 — Croatian diachronic word-embeddings paper accepted to the Slavic NLP workshop @ ACL 2025 [paper]
- May 2025 — REVS accepted to Findings of ACL 2025 [paper & code]
- Apr 2025 — MIB, a Mechanistic Interpretability Benchmark, released and accepted to ICML 2025 [paper]
Selected publications
See Google Scholar for everything else.