| 1 | Haoyu ZhangWhose Refusal Is It? The Unmeasured Contribution of Black-Box Multimodal Guardrails · 2026-08-09 1stDecoy Images Amplify Caption-Mediated Defenses Against Encoded Jailbreaks · 2026-08-02 1stOverloading Large Vision-Language Models for Jailbreaking · 2026-07-03 1st | 9.5 | 3 | 3 | 2026-08-09 |
| 2 | Han WangAdvancing Relevance Measurement with Vision-Language Models for Web-Scale Search · 2026-08-03 1stMonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Language Mo · 2026-03-30 1st | 7.0 | 2 | 2 | 2026-08-03 |
| 3 | Yang YangTrace, Verify, and Correct: A Training-Free Framework for Spatial Reasoning in Multimodal LLMs · 2026-08-05 1stStage-Transition Dense Reward Modeling for Reinforcement Learning · 2026-06-30 1st | 7.0 | 2 | 2 | 2026-08-05 |
| 4 | Abrar AlotaibiAdversarial Diffusion Across Modalities: A Fusion Survey of Attacks, Defenses, and Evaluation fo · 2026-06-25 1stA Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation · 2026-06-24 1st | 6.5 | 2 | 2 | 2026-06-25 |
| 5 | Subramanyam SahooPessimism's Paradox: Conservative Offline Training Amplifies Reward Hacking During Online Adapta · 2026-06-29 1stLinear Probes Detect Task Format, Not Reasoning Mode in Language Model Hidden States · 2026-06-01 1st | 6.5 | 2 | 2 | 2026-06-29 |
| 6 | David Demitri AfricaItem Response Theory for AI Safety · 2026-08-05Prefill Awareness in Large Language Models · 2026-06-10Consistency Training Can Entrench Misalignment · 2026-06-02 1st | 6.5 | 3 | 1 | 2026-08-05 |
| 7 | Joachim SchaefferStealing Reasoning Traces from Proprietary LLM APIs · 2026-08-10CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs · 2026-06-09 1stAttack Selection in Agentic AI Control Evaluations Meaningfully Decreases Safety · 2026-06-03 | 6.5 | 3 | 1 | 2026-08-10 |
| 8 | Kai ChenGPT-Red: Automated Red Teaming via Self-Play at Scale · 2026-07-28AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities · 2026-07-15 1stDoubtProbe: Black-Box Jailbreak Defense via Structural Verification and Semantic Auditing · 2026-06-15 | 6.5 | 3 | 1 | 2026-07-28 |
| 9 | Varad VishwarupeThe Evaluation Differential: When Frontier AI Models Recognise They Are Being Tested · 2026-05-12 1stNeurIPS Should Require Reproducibility Standards for Frontier AI Safety Claims · 2026-05-05 1st | 6.0 | 2 | 2 | 2026-05-12 |
| 10 | Mary PhuongGDM AI Control Roadmap · 2026-07-13 1stMulti-Agent AI Control: Distributed Attacks Hamper Per-Instance Monitors · 2026-07-08Bootstrapped Monitoring: Leveraging Transparent Reasoning to Oversee Stronger AI Agents · 2026-06-10 | 6.0 | 3 | 1 | 2026-07-13 |
| 11 | Yan WangYuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety · 2026-06-26Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety · 2026-06-23ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails · 2026-05-29 1st | 6.0 | 3 | 1 | 2026-06-26 |
| 12 | Fei ShenWho Bridges Safety? Identifying and Targeting Cross-Lingual Shared Safety Pathways · 2026-08-10No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks · 2026-08-02Moving the Safety Barrier: Dynamic Routing Adaptive Alignment Against White-Box Attacks · 2026-08-02 | 6.0 | 4 | 0 | 2026-08-10 |
| 13 | Tat-Seng ChuaWho Bridges Safety? Identifying and Targeting Cross-Lingual Shared Safety Pathways · 2026-08-10No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks · 2026-08-02Moving the Safety Barrier: Dynamic Routing Adaptive Alignment Against White-Box Attacks · 2026-08-02 | 6.0 | 4 | 0 | 2026-08-10 |
| 14 | Tianhang ZhengProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions · 2026-08-11DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection · 2026-07-22An Early Warning of Emerging Biosecurity Risks in Frontier LLMs · 2026-07-20 | 6.0 | 4 | 0 | 2026-08-11 |
| 15 | Yang LiuReasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces · 2026-08-12Breaking Customized LLMs for Coding: Automated Red Teaming for Instruction Backdoor Attacks · 2026-08-06Open Models, Open Risks: Measuring Unsafe Generation in Text-to-Image Models In the Wild · 2026-07-08 | 6.0 | 4 | 0 | 2026-08-12 |
| 16 | Ads DawsonStealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents · 2026-07-28 1stScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents · 2026-07-08 | 5.5 | 2 | 1 | 2026-07-28 |
| 17 | Alexander PanfilovStealing Reasoning Traces from Proprietary LLM APIs · 2026-08-10 1stCIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs · 2026-06-09 | 5.5 | 2 | 1 | 2026-08-10 |
| 18 | Axel HjmarkMeasuring Reward-Seeking via Contrastive Belief Updates · 2026-07-21 1stTraining Deliberative Monitors for Black-Box Scheming Detection · 2026-05-28 | 5.5 | 2 | 1 | 2026-07-21 |
| 19 | Harry MayneValue Leakage: An LLM's Answers Are Silently Shaped by Its Own Values · 2026-07-15Negation Neglect: When models fail to learn negations in training · 2026-05-13 1st | 5.5 | 2 | 1 | 2026-07-15 |
| 20 | Ilias ChalkidisTemplated or fully synthetic? Prompt construction as a confound in measuring LLM political stanc · 2026-08-11 1stBrainrot: Deskilling and Addiction are Overlooked AI Risks · 2026-05-05 1st | 5.5 | 2 | 1 | 2026-08-11 |
| 21 | Jing ShaoAn Early Warning of Emerging Biosecurity Risks in Frontier LLMs · 2026-07-20Mitigating Scaffolding Collapse in Socratic Tutors via Representation Alignment · 2026-06-15 1st | 5.5 | 2 | 1 | 2026-07-20 |
| 22 | Junyeong ParkEduZone: A Framework for Evaluating LLM Safety for K-12 Students and Teachers · 2026-08-03 1stPluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Relia · 2026-07-07 | 5.5 | 2 | 1 | 2026-08-03 |
| 23 | Ming LiHow Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Re · 2026-08-10 1stAgentX: Towards Agent-Driven Self-Iteration of Industrial Recommender Systems · 2026-06-25 | 5.5 | 2 | 1 | 2026-08-10 |
| 24 | Nenad TomasevAI Value Alignment for Evolving Social Norms · 2026-07-20 1stPositive Alignment: Artificial Intelligence for Human Flourishing · 2026-05-11 | 5.5 | 2 | 1 | 2026-07-20 |
| 25 | Shangze LiNo Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks · 2026-08-02Moving the Safety Barrier: Dynamic Routing Adaptive Alignment Against White-Box Attacks · 2026-08-02 1st | 5.5 | 2 | 1 | 2026-08-02 |
| 26 | Shiji ZhaoHiRoute: Hierarchical Routed Prompt Tuning for Safety Alignment of Large Language Models · 2026-08-13A Multimodal Automatic Redteaming Evaluation based on Atomic Jailbreak Strategy Decoupling and C · 2026-08-03 1st | 5.5 | 2 | 1 | 2026-08-13 |
| 27 | Shuyi MiaoWho Bridges Safety? Identifying and Targeting Cross-Lingual Shared Safety Pathways · 2026-08-10 1stOne Anchor for All: Unified Multilingual and Multimodal Safety Alignment for LVLMs · 2026-07-30 | 5.5 | 2 | 1 | 2026-08-10 |
| 28 | Simiao XieNo Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks · 2026-08-02 1stMoving the Safety Barrier: Dynamic Routing Adaptive Alignment Against White-Box Attacks · 2026-08-02 | 5.5 | 2 | 1 | 2026-08-02 |
| 29 | Viktor MoskvoretskiiSynthetic Persona Pretraining: Alignment from Token Zero · 2026-08-13Tracing Persona Vectors Through LLM Pretraining · 2026-05-13 1st | 5.5 | 2 | 1 | 2026-08-13 |
| 30 | Weiwei QiDARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection · 2026-07-22 1stDataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment · 2026-07-16 | 5.5 | 2 | 1 | 2026-07-22 |
| 31 | Xucheng YuUnderstanding Content Moderation in Large Language Models through Restricted Books: From Refusal · 2026-08-12 1stMonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Language Mo · 2026-03-30 | 5.5 | 2 | 1 | 2026-08-12 |
| 32 | Youting WangSafety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks · 2026-07-30 1stSelf-Commitment Latency: A Reward-Free Probe for Prompted Implicit Hacking · 2026-06-04 | 5.5 | 2 | 1 | 2026-07-30 |
| 33 | Yuchen ChenBreaking Customized LLMs for Coding: Automated Red Teaming for Instruction Backdoor Attacks · 2026-08-06 1stExecution-Grounded Security Testing for Coding Agents in Software Engineering Pipelines · 2026-06-01 | 5.5 | 2 | 1 | 2026-08-06 |
| 34 | Zefeng WuDARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection · 2026-07-22DataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment · 2026-07-16 1st | 5.5 | 2 | 1 | 2026-07-22 |
| 35 | Hui XueOyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models · 2026-07-03Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety · 2026-06-26Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety · 2026-06-23 | 5.5 | 4 | 0 | 2026-07-03 |
| 36 | Shikai QiuYuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety · 2026-06-26Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety · 2026-06-23 1st | 5.0 | 2 | 1 | 2026-06-26 |
| 37 | Ting MaYuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety · 2026-06-26 1stYuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety · 2026-06-23 | 5.0 | 2 | 1 | 2026-06-26 |
| 38 | Ashton AndersonSynthetic Persona Pretraining: Alignment from Token Zero · 2026-08-13Studying People to Study AI: Expert Perspectives on the Epistemic Fit and Barriers of Human Rese · 2026-08-06Grounded Chess Reasoning in Language Models via Master Distillation · 2026-03-20 | 5.0 | 3 | 0 | 2026-08-13 |
| 39 | Bin LiuYesterday's Shield, Today's Spear: A Self-Evolving Safety Guardrail in Production · 2026-08-09Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety · 2026-06-26Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety · 2026-06-23 | 5.0 | 3 | 0 | 2026-08-09 |
| 40 | Chaochao LuDARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection · 2026-07-22An Early Warning of Emerging Biosecurity Risks in Frontier LLMs · 2026-07-20DataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment · 2026-07-16 | 5.0 | 3 | 0 | 2026-07-22 |
| 41 | Chuancheng ShiNo Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks · 2026-08-02Moving the Safety Barrier: Dynamic Routing Adaptive Alignment Against White-Box Attacks · 2026-08-02One Anchor for All: Unified Multilingual and Multimodal Safety Alignment for LVLMs · 2026-07-30 | 5.0 | 3 | 0 | 2026-08-02 |
| 42 | Guanghui WangReconcile Once, Write Anytime: A Trust-Tiered Librarian and a Multi-Agent Writer for Drift-Free, · 2026-08-13Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety · 2026-06-26Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety · 2026-06-23 | 5.0 | 3 | 0 | 2026-08-13 |
| 43 | Jonas GeipingStealing Reasoning Traces from Proprietary LLM APIs · 2026-08-10CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs · 2026-06-09Models That Know How Evaluations Are Designed Score Safer · 2026-05-27 | 5.0 | 3 | 0 | 2026-08-10 |
| 44 | Kui RenDARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection · 2026-07-22DataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment · 2026-07-16ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails · 2026-05-29 | 5.0 | 3 | 0 | 2026-07-22 |
| 45 | Liang HeDARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection · 2026-07-22DataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment · 2026-07-16Caring Without Feeling: Affective Dynamics as the Control Layer of Human-AI Agent Collaboration · 2026-05-08 | 5.0 | 3 | 0 | 2026-07-22 |
| 46 | Qian WangMMAligner: Safeguarding Multimodal Large Language Models through Representation Calibration · 2026-08-06Enhancing Multimodal In-Context Learning via Inductive-Deductive Reasoning · 2026-05-04Are Dilemmas and Conflicts in LLM Alignment Solvable? A View from Priority Graph · 2026-03-16 | 5.0 | 3 | 0 | 2026-08-06 |
| 47 | Xia HuOpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution · 2026-08-01Do LLMs Know Their Vulnerable Scenarios? · 2026-07-26An Early Warning of Emerging Biosecurity Risks in Frontier LLMs · 2026-07-20 | 5.0 | 3 | 0 | 2026-08-01 |
| 48 | Zhan QinDARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection · 2026-07-22DataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment · 2026-07-16Adaptive and Explicit safe: Triggering Latent Safety Awareness in Large Reasoning Models · 2026-06-15 | 5.0 | 3 | 0 | 2026-07-22 |
| 49 | Zonghao YingSafeFlow: Semantic Information-Flow Control for Blocking Malicious Propagation in Multi-Agent Sy · 2026-07-28Dynamic Defense Profiling Enables Cognitive Jailbreak of Text-to-Image Models · 2026-07-20Securing the AI Agent: A Unified Framework for Multi-Layer Agent Red Teaming · 2026-06-30 | 5.0 | 3 | 0 | 2026-07-28 |
| 50 | Georgina CosmaSUPREME: A Multi-GPU Framework for Reproducible Image Unlearning Method Evaluation · 2026-05-29RULER: Representation-Level Verification of Machine Unlearning · 2026-05-26 1st | 4.5 | 2 | 1 | 2026-05-29 |
| 51 | Jiaheng WeiGeoFaith: A Spatio-Temporal Dual View of Faithful Chain-of-Thought · 2026-05-26Rethinking Federated Unlearning via the Lens of Memorization · 2026-05-23 1st | 4.5 | 2 | 1 | 2026-05-26 |
| 52 | Xuyang ZhongA Full-Pipeline Framework for Evaluating Membership Inference Attacks in Machine Learning · 2026-05-28DualOptim+: Bridging Shared and Decoupled Optimizer States for Better Machine Unlearning in Larg · 2026-05-20 1st | 4.5 | 2 | 1 | 2026-05-28 |
| 53 | Aman ChadhaPessimism's Paradox: Conservative Offline Training Amplifies Reward Hacking During Online Adapta · 2026-06-29Linear Probes Detect Task Format, Not Reasoning Mode in Language Model Hidden States · 2026-06-01MAAT: Multi-phase Adapter-Aware Targeted Unlearning · 2026-05-28 | 4.5 | 3 | 0 | 2026-06-29 |
| 54 | Bingyu ZhuYuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety · 2026-06-26Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety · 2026-06-23ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails · 2026-05-29 | 4.5 | 3 | 0 | 2026-06-26 |
| 55 | Bo LiPolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails · 2026-07-07MAStrike: Shapley-Guided Collusive Red-Teaming on Multi-Agent Systems · 2026-06-11Are Dilemmas and Conflicts in LLM Alignment Solvable? A View from Priority Graph · 2026-03-16 | 4.5 | 3 | 0 | 2026-07-07 |
| 56 | Fazl BarezPretraining Curricula Enable Selective Fine-tuning · 2026-07-06The Capability Frontier: Benchmarks Miss 82% of Model Performance · 2026-06-25Position: Don't Just "Fix it in Post": A Science of AI Must Study Training Dynamics · 2026-06-03 | 4.5 | 3 | 0 | 2026-07-06 |
| 57 | Jing WangYuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety · 2026-06-26AgentX: Towards Agent-Driven Self-Iteration of Industrial Recommender Systems · 2026-06-25Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety · 2026-06-23 | 4.5 | 3 | 0 | 2026-06-26 |
| 58 | Longtao HuangYuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety · 2026-06-26Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety · 2026-06-23ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails · 2026-05-29 | 4.5 | 3 | 0 | 2026-06-26 |
| 59 | Vinija JainPessimism's Paradox: Conservative Offline Training Amplifies Reward Hacking During Online Adapta · 2026-06-29Linear Probes Detect Task Format, Not Reasoning Mode in Language Model Hidden States · 2026-06-01MAAT: Multi-phase Adapter-Aware Targeted Unlearning · 2026-05-28 | 4.5 | 3 | 0 | 2026-06-29 |
| 60 | Wei WangDT-Guard: Intent-Driven Reasoning-Active Training for Reasoning-Free LLM Safety Guardrail · 2026-07-07Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety · 2026-06-26Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety · 2026-06-23 | 4.5 | 3 | 0 | 2026-07-07 |
| 61 | Adam GleaveAI Security Leaderboard: Methodology, Results and Minimal Standard · 2026-08-04Scaling Trends for Lie Detector Oversight in Preference Learning · 2026-07-02 | 4.0 | 2 | 0 | 2026-08-04 |
| 62 | Alice OhEduZone: A Framework for Evaluating LLM Safety for K-12 Students and Teachers · 2026-08-03Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Relia · 2026-07-07 | 4.0 | 2 | 0 | 2026-08-03 |
| 63 | Chao ShenMMAligner: Safeguarding Multimodal Large Language Models through Representation Calibration · 2026-08-06Alignment Is Local: A Paired Diagnostic for GUI Agents under User-Side Persuasion · 2026-07-31 | 4.0 | 2 | 0 | 2026-08-06 |
| 64 | Chunrong FangBreaking Customized LLMs for Coding: Automated Red Teaming for Instruction Backdoor Attacks · 2026-08-06Execution-Grounded Security Testing for Coding Agents in Software Engineering Pipelines · 2026-06-01 | 4.0 | 2 | 0 | 2026-08-06 |
| 65 | David SchmotzStealing Reasoning Traces from Proprietary LLM APIs · 2026-08-10ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D · 2026-07-21 | 4.0 | 2 | 0 | 2026-08-10 |
| 66 | Di WangProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions · 2026-08-11Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs · 2026-08-10 | 4.0 | 2 | 0 | 2026-08-11 |
| 67 | Dingyan ShangSafety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks · 2026-07-30Self-Commitment Latency: A Reward-Free Probe for Prompted Implicit Hacking · 2026-06-04 | 4.0 | 2 | 0 | 2026-07-30 |
| 68 | Dongbin NaWhen Are Reasoning-Based Guardrails Not Efficient? ResponseGuard: A Fast Vision-Language Guard f · 2026-07-23 1stDo Safety Guardrails Need to Reason? LeanGuard: A Fast and Light Approach for Robust Moderation · 2026-06-25 1st | 4.0 | 2 | 0 | 2026-07-23 |
| 69 | Feng ChenREDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems · 2026-08-11Stress Testing Concept Erasure with Large Language Model Agents · 2026-07-20 | 4.0 | 2 | 0 | 2026-08-11 |
| 70 | Hangtao ZhangTYPO: Instruction-Dense Visual Jailbreaks against Commercial Closed-Source Image-Generation Mode · 2026-07-27PVDetector: Detecting Prompt Injection Attacks on Purpose-Specific LLM Agents through Policy-Vio · 2026-07-14 | 4.0 | 2 | 0 | 2026-07-27 |
| 71 | Jan DubinskiValue Leakage: An LLM's Answers Are Silently Shaped by Its Own Values · 2026-07-15Negation Neglect: When models fail to learn negations in training · 2026-05-13 | 4.0 | 2 | 0 | 2026-07-15 |
| 72 | Jannik BrinkmannSynthetic Persona Pretraining: Alignment from Token Zero · 2026-08-13Mood Matters: How Syntactic Sensitivity Undermines Safety Alignment · 2026-08-05 | 4.0 | 2 | 0 | 2026-08-13 |
| 73 | Jiawei ChenTrace, Verify, and Correct: A Training-Free Framework for Spatial Reasoning in Multimodal LLMs · 2026-08-05The Verification Horizon: No Silver Bullet for Coding Agent Rewards · 2026-06-24 | 4.0 | 2 | 0 | 2026-08-05 |
| 74 | Jie ZhangBenchmarking Cyberattack Detection in Electric Vehicle Charging Infrastructure with Benign User · 2026-08-11An Early Warning of Emerging Biosecurity Risks in Frontier LLMs · 2026-07-20 | 4.0 | 2 | 0 | 2026-08-11 |
| 75 | Jing LiWhen Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Sk · 2026-08-09ASRU: Activation Steering Meets Reinforcement Unlearning for Multimodal Large Language Models · 2026-05-15 | 4.0 | 2 | 0 | 2026-08-09 |
| 76 | Jinqiu YangChecked-In Secret Detection: Strings Are All You Need · 2026-08-05IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests · 2026-07-22 | 4.0 | 2 | 0 | 2026-08-05 |
| 77 | Joe BentonDiffuse AI Control on Fuzzy Tasks · 2026-06-08Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning · 2026-05-22Removing Sandbagging in LLMs by Training with Weak Supervision · 2026-04-23 | 4.0 | 3 | 0 | 2026-06-08 |
| 78 | Junkai ChenExploring and Bridging Knowledge Holes in Unlearned Multimodal Large Language Models · 2026-08-03Toward Fine-Grained Forgetting:Attribute Unlearning for Multimodal Large Language Models · 2026-08-02 | 4.0 | 2 | 0 | 2026-08-03 |
| 79 | Juntao DaiA Blind Spot in Alignment: Quantifying Biosecurity Risks in Large Language Models · 2026-08-03When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agents · 2026-06-18 | 4.0 | 2 | 0 | 2026-08-03 |
| 80 | Li ZengTYPO: Instruction-Dense Visual Jailbreaks against Commercial Closed-Source Image-Generation Mode · 2026-07-27PVDetector: Detecting Prompt Injection Attacks on Purpose-Specific LLM Agents through Policy-Vio · 2026-07-14 | 4.0 | 2 | 0 | 2026-07-27 |