Zihe Yan

PhD Researcher · Shanghai Jiao Tong University

Zihe Yan 闫子赫

I study the safety of multimodal AI agents, with a focus on GUI attacks, model interpretability, and aligning agent behavior with user intent.

  • Multimodal and GUI agent safety
  • LLM interpretability and representation engineering
  • Agent harness and access control
  • User-intent alignment and adversarial evaluation

Published

LaSM layer-27 attention heatmaps comparing scaling layers 21 to 26, no scaling, and scaling layers 7 to 18.
link to full paper

CVPR 2026 · First author

LaSM: Layer-wise Scaling Mechanism for Defending Pop-up Attack on GUI Agents

Zihe Yan, Jiaping Gui, Zhuosheng Zhang, Gongshen Liu

Defending GUI agents against pop-up attacks through layer-wise scaling.

Backdoor risk assessment architecture with metadata inputs, four evaluation modules, risk integration, and explainable analysis.
link to full paper

AAAI 2026 · Co-first author

An LLM-based Quantitative Framework for Evaluating High-Stealthy Backdoor Risks in OSS Supply Chains

Zihe Yan, Kai Luo, Haoyu Yang, Yang Yu, Zhuosheng Zhang, Guancheng Li

Using LLMs to quantitatively assess stealthy backdoor risks in open-source software supply chains.

EVA framework showing offline semantic attack discovery, a rule library, and online deployment.
link to full paper

ACL 2026 · Fourth author

EVA: Evolving Semantic Adversaries for Red-Teaming GUI Agents Against Environmental Injection Attacks

Yijie Lu, Manman Zhao, Tianjie Ju, Zihe Yan, Xinbei Ma, Yuan Guo, Daizong Ding, Gongshen Liu, Zhuosheng Zhang

Red-teaming GUI agents with evolving semantic adversaries and environmental injection attacks.

Implementation process of the proposed secure identity resolution scheme, showing initialization and secure identity resolution phases with policy matching, data queries, and traceability.
link to full paper

IEEE Internet of Things Journal 2024 · Second author

Attribute-Based Access Control Scheme for Secure Identity Resolution in Prognostics and Health Management

Yunhua He, Zihe Yan, Tingli Yuan

Attribute-based access control for secure identity resolution in industrial systems.

Submitted

Submitted to ICASSP 2027 · First author

DPIM: A Dual-probe Intervention Mechanism For Omni-LLM Security Enhancement.

Submitted to EACL 2026 · First author

Representation Engineering in Vision Language Models: A survey

Submitted to ICLR 2027 · Second author

RASER: Recoverability-Aware Selective Escalation Router for Multi-Hop Question Answering

Submitted to ICLR 2027 · First author

Aligning with the Wrong Intent: How User-Intent Ambiguity Enables Semantic Attacks on MLLM Agents

In Preparation

Manuscript in preparation

From Prompt to Permission: Securing Personal Agents with Access Control Policies

I am a PhD student in Cyberspace Security at Shanghai Jiao Tong University, advised by Zhuosheng Zhang and Jiaping Gui. My research connects multimodal agent safety with model interpretability and practical security controls.

My earlier work explored industrial internet security, identity resolution, cryptography, and blockchain. My current research focuses on MLLM agent security, with an emphasis on both internal representation-level interventions and external safety harnesses for improving agent robustness and behavioral safety.

I received national third prizes in both the 7th and 8th National College Student Cryptography Competitions.

Sep 2024–present · Expected Jun 2028

Shanghai Jiao Tong University

PhD in Cyberspace Security

Sep 2021–Jun 2024

North China University of Technology

Master’s in Cyberspace Security

Sep 2017–Jun 2021

Nanjing Tech University

Bachelor’s in Process Equipment and Control Engineering

2026 · Research internship

Huawei 2012 Lab

Huawei 2012 Lab

Research on agent intent analysis, behavioral safety alignment, and representation engineering for multimodal model safety.

2025 · Course development

Huawei Ascend

Huawei Ascend

Contributed lessons on MindStudio and large language model evaluation to Dive into LLMs (★ 52.9k), a course covering the LLM development lifecycle.

2022 · Research Internship

China Academy of Industrial Internet

China Academy of Industrial Internet

Worked on industrial cybersecurity research, contributed to security standardization, and conducted field studies at manufacturing enterprises to understand practical security requirements and challenges in industrial systems.

2019 · Exchange Program

RWTH Aachen University

RWTH Aachen University

Studied industrial robotics, augmented reality systems, and intelligent manufacturing through coursework and project-based learning.

Based in Shanghai. For research discussions and collaboration, please get in touch by email.