Egor Zverev

Hi! My name is Egor and I am a PhD student at ISTA, Austria, working under the supervision of Christoph Lampert. I am also a member of the ELLIS PhD Program co-supervised by Florian Tramèr from ETH Zürich. Since June 2026, I have been a research intern at Microsoft Research Cambridge, working with the Project Roma team on secure-by-design web agents, and I also advise several master's students on AI security in robotics through the ETH Robotics Club.

I am broadly interested in AI Safety and Security, with a particular focus on improving LLM Security through architectures and building provable safeguards for AI agents. In my previous work I have formalized the problem of instruction-data separation (i.e., what it means for the model to separate executable instructions from non-executable data) and proposed a method to increase such separation through architectural changes. More recently, I have been working on provable security guarantees for web agents that need to handle untrusted content.

Before coming to ISTA, I got my B.S.(Hons) in Applied Math and CS from the Yandex Department of Data Analysis at the Moscow Institute of Physics and Technology, where I also taught Python and stats. You can find my full CV here.

I also enjoy doing creative and community-driven gigs. I was the lead organizer for the Foundations of LLM Security Workshop @EurIPS'25 and the LLM Safety and Security ELLIS UnConference'25, and I am now co-organizing the 2nd iteration of the Foundations of LLM Security Workshop @NeurIPS'26. I co-created one of the AI art booths for ISTA Summer Campus, and in my spare time, I write poetry and tinker with Arduino.

I am always open to new connections, professional and otherwise, feel free to send me an email at [first_name].[last_name]@ist.ac.at. Let's chat!

Egor Zverev

Publications

Untrusted Content Masking for Web Agents with Security Guarantees

Kristina Nikolić*, Egor Zverev*, Javier Rando, Matthew Jagielski, Edoardo Debenedetti, Florian Tramèr (*equal contribution)
ICML Workshop on Agents in the Wild (AIWILD), 2026 (Spotlight)

ASIDE: Architectural Separation of Instructions and Data in Language Models

Egor Zverev, Evgenii Kortukov, Alexander Panfilov, Alexandra Volkova, Soroush Tabesh, Sebastian Lapuschkin, Wojciech Samek, Christoph H. Lampert
ICLR 2026
ICLR 2025 Workshop on Building Trust in Language Models and Applications (Oral)

Can LLMs Separate Instructions From Data? And What Do We Even Mean By That?

Egor Zverev, Sahar Abdelnabi, Soroush Tabesh, Mario Fritz, Christoph H. Lampert
ICLR 2025

LLMail-Inject: A Dataset from a Realistic Adaptive Prompt Injection Challenge

Sahar Abdelnabi, Aideen Fay, Ahmed Salem, Egor Zverev, Kai-Chieh Liao, ... Andrew Paverd, Giovanni Cherubin
EurIPS 2025 Salon des Refusés (Accepted by AC/SAC at NeurIPS'25, rejected by PC due to venue capacity constraints)

News

July 2026
Gave an invited talk at Imperial College London on our new paper, "Untrusted Content Masking for Web Agents with Security Guarantees".
July 2026
Excited to share that our Foundations of LLM Security Workshop got accepted to NeurIPS'26 in Paris! I am co-leading the organization. So excited about our speaker lineup: Niloofar Mireshghallah, Christian Schroeder de Witt, Reza Shokri, Somesh Jha and David Evans. If you work on LLM security, consider submitting a paper!
June 2026
Our new paper "Untrusted Content Masking for Web Agents with Security Guarantees" (joint work with Kristina Nikolić, equal contribution) got a Spotlight at the ICML Workshop on Agents in the Wild! Check it out here.
June 2026
I started a research internship at Microsoft Research Cambridge, working with Santiago Zanella-Béguelin on secure-by-design web agents.
March 2026
I went on a "Science Tour", giving talks on Untrusted Content Masking and ASIDE at ELLIS Institute Tübingen, Saarland University, University of Luxembourg, University of Liechtenstein, and CISPA, all within 5 days.
March 2026
Became a Research Advisor at the ETH Robotics Club, supervising Andrei Baroian, Noé Macé, and Gustave Charles Saigne on a project exploring AI security vulnerabilities in state-of-the-art robotics models.
February 2026
I am attending IASEAI'26 in Paris. I will give a talk about ASIDE and be a panelist in the follow-up session. Exciting!
January 2026
ASIDE got accepted to ICLR 2026!
October 2025
I am co-organizing the Foundations of LLM Security Workshop @EurIPS'25! Call for talks open until Oct 22. Proud of the speakers we have at the event: Ilia Shumailov, Santiago Zanella-Béguelin, Kathrin Grosse and Verena Rieser.
October 2025
I am co-organizing the LLM Safety and Security ELLIS UnConference'25 workshop! We have really nice speakers: Isabel Valera, Pepa Atanasova and Qiongxiu Li.
September 2025
I have co-affiliated with Prof. Florian Tramèr (ETHZ) through ELLIS PhD Program! I am visiting ETHZ from September to December to work on LLM Agents Security.
September 2025
Gave a talk on instruction-data separation at IBM T.J. Watson Research Lab. Thank you Rosario Uceda-Sosa for hosting me!
August 2025
Attended CISPA - ELLIS - Summer School 2025 on Trustworthy AI in Saarbrücken, Germany.
June 2025
Our paper on LLMail-Inject competition is out. Check it out here. Had a fun time collaborating with Microsoft on this!
May 2025
Attended ICLR 2025 in Singapore to present SEP and ASIDE papers. Gave a talk about ASIDE at the BuildTrust workshop.
March 2025
Our paper "ASIDE: Architectural Separation of Instructions and Data in Language Models" was accepted to ICLR 2025 BuildTrust workshop for an oral presentation!
March 2025
Gave an invited talk at ETH about ASIDE. Grateful to Florian Tramèr for hosting me!
January 2025
Our paper "Can LLMs Separate Instructions From Data? And What Do We Even Mean By That?" was accepted to ICLR 2025!
October 2024
I am co-organizing LLMail-Inject competition with Microsoft. Check it out here.