Juniper Ventures · internal · sourcing engine v0

Sourcing engine v0 — arXiv AI-safety candidate feed.

The first running slice of the sourcing engine: it pulls recent AI-safety / assurance papers from arXiv, aggregates by author, and ranks by first-authorship (the builder — the "researcher whose paper is the product," which is Juniper's thesis). This is a raw candidate pool to human-filter for founder potential, not a finished list.

generated 2026-08-14 633 safety papers · 160d window 184 authors ≥2 papers
v0 caveats (read before acting): (1) Name-only dedup — common names (e.g. "Yang Liu") merge multiple different people and rank falsely high; open the papers to disambiguate. (2) Many top authors are at frontier labs, not founder candidates — this feeds the human "founder-potential" filter, it doesn't replace it. (3) arXiv only — v1 adds affiliation disambiguation (OpenReview API) + Apart / LessWrong / 80k feeds (see the talent map). Score = 2.5×first-author + 1×other papers + recency.

Ranked candidates (80 of 184)

#Researcher · recent AI-safety papersscorepapers1st-authlatest
1Haoyu ZhangWhose Refusal Is It? The Unmeasured Contribution of Black-Box Multimodal Guardrails · 2026-08-09 1stDecoy Images Amplify Caption-Mediated Defenses Against Encoded Jailbreaks · 2026-08-02 1stOverloading Large Vision-Language Models for Jailbreaking · 2026-07-03 1st9.5332026-08-09
2Han WangAdvancing Relevance Measurement with Vision-Language Models for Web-Scale Search · 2026-08-03 1stMonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Language Mo · 2026-03-30 1st7.0222026-08-03
3Yang YangTrace, Verify, and Correct: A Training-Free Framework for Spatial Reasoning in Multimodal LLMs · 2026-08-05 1stStage-Transition Dense Reward Modeling for Reinforcement Learning · 2026-06-30 1st7.0222026-08-05
4Abrar AlotaibiAdversarial Diffusion Across Modalities: A Fusion Survey of Attacks, Defenses, and Evaluation fo · 2026-06-25 1stA Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation · 2026-06-24 1st6.5222026-06-25
5Subramanyam SahooPessimism's Paradox: Conservative Offline Training Amplifies Reward Hacking During Online Adapta · 2026-06-29 1stLinear Probes Detect Task Format, Not Reasoning Mode in Language Model Hidden States · 2026-06-01 1st6.5222026-06-29
6David Demitri AfricaItem Response Theory for AI Safety · 2026-08-05Prefill Awareness in Large Language Models · 2026-06-10Consistency Training Can Entrench Misalignment · 2026-06-02 1st6.5312026-08-05
7Joachim SchaefferStealing Reasoning Traces from Proprietary LLM APIs · 2026-08-10CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs · 2026-06-09 1stAttack Selection in Agentic AI Control Evaluations Meaningfully Decreases Safety · 2026-06-036.5312026-08-10
8Kai ChenGPT-Red: Automated Red Teaming via Self-Play at Scale · 2026-07-28AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities · 2026-07-15 1stDoubtProbe: Black-Box Jailbreak Defense via Structural Verification and Semantic Auditing · 2026-06-156.5312026-07-28
9Varad VishwarupeThe Evaluation Differential: When Frontier AI Models Recognise They Are Being Tested · 2026-05-12 1stNeurIPS Should Require Reproducibility Standards for Frontier AI Safety Claims · 2026-05-05 1st6.0222026-05-12
10Mary PhuongGDM AI Control Roadmap · 2026-07-13 1stMulti-Agent AI Control: Distributed Attacks Hamper Per-Instance Monitors · 2026-07-08Bootstrapped Monitoring: Leveraging Transparent Reasoning to Oversee Stronger AI Agents · 2026-06-106.0312026-07-13
11Yan WangYuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety · 2026-06-26Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety · 2026-06-23ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails · 2026-05-29 1st6.0312026-06-26
12Fei ShenWho Bridges Safety? Identifying and Targeting Cross-Lingual Shared Safety Pathways · 2026-08-10No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks · 2026-08-02Moving the Safety Barrier: Dynamic Routing Adaptive Alignment Against White-Box Attacks · 2026-08-026.0402026-08-10
13Tat-Seng ChuaWho Bridges Safety? Identifying and Targeting Cross-Lingual Shared Safety Pathways · 2026-08-10No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks · 2026-08-02Moving the Safety Barrier: Dynamic Routing Adaptive Alignment Against White-Box Attacks · 2026-08-026.0402026-08-10
14Tianhang ZhengProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions · 2026-08-11DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection · 2026-07-22An Early Warning of Emerging Biosecurity Risks in Frontier LLMs · 2026-07-206.0402026-08-11
15Yang LiuReasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces · 2026-08-12Breaking Customized LLMs for Coding: Automated Red Teaming for Instruction Backdoor Attacks · 2026-08-06Open Models, Open Risks: Measuring Unsafe Generation in Text-to-Image Models In the Wild · 2026-07-086.0402026-08-12
16Ads DawsonStealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents · 2026-07-28 1stScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents · 2026-07-085.5212026-07-28
17Alexander PanfilovStealing Reasoning Traces from Proprietary LLM APIs · 2026-08-10 1stCIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs · 2026-06-095.5212026-08-10
18Axel HjmarkMeasuring Reward-Seeking via Contrastive Belief Updates · 2026-07-21 1stTraining Deliberative Monitors for Black-Box Scheming Detection · 2026-05-285.5212026-07-21
19Harry MayneValue Leakage: An LLM's Answers Are Silently Shaped by Its Own Values · 2026-07-15Negation Neglect: When models fail to learn negations in training · 2026-05-13 1st5.5212026-07-15
20Ilias ChalkidisTemplated or fully synthetic? Prompt construction as a confound in measuring LLM political stanc · 2026-08-11 1stBrainrot: Deskilling and Addiction are Overlooked AI Risks · 2026-05-05 1st5.5212026-08-11
21Jing ShaoAn Early Warning of Emerging Biosecurity Risks in Frontier LLMs · 2026-07-20Mitigating Scaffolding Collapse in Socratic Tutors via Representation Alignment · 2026-06-15 1st5.5212026-07-20
22Junyeong ParkEduZone: A Framework for Evaluating LLM Safety for K-12 Students and Teachers · 2026-08-03 1stPluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Relia · 2026-07-075.5212026-08-03
23Ming LiHow Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Re · 2026-08-10 1stAgentX: Towards Agent-Driven Self-Iteration of Industrial Recommender Systems · 2026-06-255.5212026-08-10
24Nenad TomasevAI Value Alignment for Evolving Social Norms · 2026-07-20 1stPositive Alignment: Artificial Intelligence for Human Flourishing · 2026-05-115.5212026-07-20
25Shangze LiNo Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks · 2026-08-02Moving the Safety Barrier: Dynamic Routing Adaptive Alignment Against White-Box Attacks · 2026-08-02 1st5.5212026-08-02
26Shiji ZhaoHiRoute: Hierarchical Routed Prompt Tuning for Safety Alignment of Large Language Models · 2026-08-13A Multimodal Automatic Redteaming Evaluation based on Atomic Jailbreak Strategy Decoupling and C · 2026-08-03 1st5.5212026-08-13
27Shuyi MiaoWho Bridges Safety? Identifying and Targeting Cross-Lingual Shared Safety Pathways · 2026-08-10 1stOne Anchor for All: Unified Multilingual and Multimodal Safety Alignment for LVLMs · 2026-07-305.5212026-08-10
28Simiao XieNo Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks · 2026-08-02 1stMoving the Safety Barrier: Dynamic Routing Adaptive Alignment Against White-Box Attacks · 2026-08-025.5212026-08-02
29Viktor MoskvoretskiiSynthetic Persona Pretraining: Alignment from Token Zero · 2026-08-13Tracing Persona Vectors Through LLM Pretraining · 2026-05-13 1st5.5212026-08-13
30Weiwei QiDARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection · 2026-07-22 1stDataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment · 2026-07-165.5212026-07-22
31Xucheng YuUnderstanding Content Moderation in Large Language Models through Restricted Books: From Refusal · 2026-08-12 1stMonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Language Mo · 2026-03-305.5212026-08-12
32Youting WangSafety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks · 2026-07-30 1stSelf-Commitment Latency: A Reward-Free Probe for Prompted Implicit Hacking · 2026-06-045.5212026-07-30
33Yuchen ChenBreaking Customized LLMs for Coding: Automated Red Teaming for Instruction Backdoor Attacks · 2026-08-06 1stExecution-Grounded Security Testing for Coding Agents in Software Engineering Pipelines · 2026-06-015.5212026-08-06
34Zefeng WuDARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection · 2026-07-22DataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment · 2026-07-16 1st5.5212026-07-22
35Hui XueOyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models · 2026-07-03Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety · 2026-06-26Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety · 2026-06-235.5402026-07-03
36Shikai QiuYuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety · 2026-06-26Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety · 2026-06-23 1st5.0212026-06-26
37Ting MaYuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety · 2026-06-26 1stYuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety · 2026-06-235.0212026-06-26
38Ashton AndersonSynthetic Persona Pretraining: Alignment from Token Zero · 2026-08-13Studying People to Study AI: Expert Perspectives on the Epistemic Fit and Barriers of Human Rese · 2026-08-06Grounded Chess Reasoning in Language Models via Master Distillation · 2026-03-205.0302026-08-13
39Bin LiuYesterday's Shield, Today's Spear: A Self-Evolving Safety Guardrail in Production · 2026-08-09Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety · 2026-06-26Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety · 2026-06-235.0302026-08-09
40Chaochao LuDARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection · 2026-07-22An Early Warning of Emerging Biosecurity Risks in Frontier LLMs · 2026-07-20DataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment · 2026-07-165.0302026-07-22
41Chuancheng ShiNo Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks · 2026-08-02Moving the Safety Barrier: Dynamic Routing Adaptive Alignment Against White-Box Attacks · 2026-08-02One Anchor for All: Unified Multilingual and Multimodal Safety Alignment for LVLMs · 2026-07-305.0302026-08-02
42Guanghui WangReconcile Once, Write Anytime: A Trust-Tiered Librarian and a Multi-Agent Writer for Drift-Free, · 2026-08-13Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety · 2026-06-26Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety · 2026-06-235.0302026-08-13
43Jonas GeipingStealing Reasoning Traces from Proprietary LLM APIs · 2026-08-10CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs · 2026-06-09Models That Know How Evaluations Are Designed Score Safer · 2026-05-275.0302026-08-10
44Kui RenDARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection · 2026-07-22DataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment · 2026-07-16ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails · 2026-05-295.0302026-07-22
45Liang HeDARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection · 2026-07-22DataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment · 2026-07-16Caring Without Feeling: Affective Dynamics as the Control Layer of Human-AI Agent Collaboration · 2026-05-085.0302026-07-22
46Qian WangMMAligner: Safeguarding Multimodal Large Language Models through Representation Calibration · 2026-08-06Enhancing Multimodal In-Context Learning via Inductive-Deductive Reasoning · 2026-05-04Are Dilemmas and Conflicts in LLM Alignment Solvable? A View from Priority Graph · 2026-03-165.0302026-08-06
47Xia HuOpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution · 2026-08-01Do LLMs Know Their Vulnerable Scenarios? · 2026-07-26An Early Warning of Emerging Biosecurity Risks in Frontier LLMs · 2026-07-205.0302026-08-01
48Zhan QinDARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection · 2026-07-22DataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment · 2026-07-16Adaptive and Explicit safe: Triggering Latent Safety Awareness in Large Reasoning Models · 2026-06-155.0302026-07-22
49Zonghao YingSafeFlow: Semantic Information-Flow Control for Blocking Malicious Propagation in Multi-Agent Sy · 2026-07-28Dynamic Defense Profiling Enables Cognitive Jailbreak of Text-to-Image Models · 2026-07-20Securing the AI Agent: A Unified Framework for Multi-Layer Agent Red Teaming · 2026-06-305.0302026-07-28
50Georgina CosmaSUPREME: A Multi-GPU Framework for Reproducible Image Unlearning Method Evaluation · 2026-05-29RULER: Representation-Level Verification of Machine Unlearning · 2026-05-26 1st4.5212026-05-29
51Jiaheng WeiGeoFaith: A Spatio-Temporal Dual View of Faithful Chain-of-Thought · 2026-05-26Rethinking Federated Unlearning via the Lens of Memorization · 2026-05-23 1st4.5212026-05-26
52Xuyang ZhongA Full-Pipeline Framework for Evaluating Membership Inference Attacks in Machine Learning · 2026-05-28DualOptim+: Bridging Shared and Decoupled Optimizer States for Better Machine Unlearning in Larg · 2026-05-20 1st4.5212026-05-28
53Aman ChadhaPessimism's Paradox: Conservative Offline Training Amplifies Reward Hacking During Online Adapta · 2026-06-29Linear Probes Detect Task Format, Not Reasoning Mode in Language Model Hidden States · 2026-06-01MAAT: Multi-phase Adapter-Aware Targeted Unlearning · 2026-05-284.5302026-06-29
54Bingyu ZhuYuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety · 2026-06-26Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety · 2026-06-23ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails · 2026-05-294.5302026-06-26
55Bo LiPolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails · 2026-07-07MAStrike: Shapley-Guided Collusive Red-Teaming on Multi-Agent Systems · 2026-06-11Are Dilemmas and Conflicts in LLM Alignment Solvable? A View from Priority Graph · 2026-03-164.5302026-07-07
56Fazl BarezPretraining Curricula Enable Selective Fine-tuning · 2026-07-06The Capability Frontier: Benchmarks Miss 82% of Model Performance · 2026-06-25Position: Don't Just "Fix it in Post": A Science of AI Must Study Training Dynamics · 2026-06-034.5302026-07-06
57Jing WangYuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety · 2026-06-26AgentX: Towards Agent-Driven Self-Iteration of Industrial Recommender Systems · 2026-06-25Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety · 2026-06-234.5302026-06-26
58Longtao HuangYuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety · 2026-06-26Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety · 2026-06-23ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails · 2026-05-294.5302026-06-26
59Vinija JainPessimism's Paradox: Conservative Offline Training Amplifies Reward Hacking During Online Adapta · 2026-06-29Linear Probes Detect Task Format, Not Reasoning Mode in Language Model Hidden States · 2026-06-01MAAT: Multi-phase Adapter-Aware Targeted Unlearning · 2026-05-284.5302026-06-29
60Wei WangDT-Guard: Intent-Driven Reasoning-Active Training for Reasoning-Free LLM Safety Guardrail · 2026-07-07Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety · 2026-06-26Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety · 2026-06-234.5302026-07-07
61Adam GleaveAI Security Leaderboard: Methodology, Results and Minimal Standard · 2026-08-04Scaling Trends for Lie Detector Oversight in Preference Learning · 2026-07-024.0202026-08-04
62Alice OhEduZone: A Framework for Evaluating LLM Safety for K-12 Students and Teachers · 2026-08-03Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Relia · 2026-07-074.0202026-08-03
63Chao ShenMMAligner: Safeguarding Multimodal Large Language Models through Representation Calibration · 2026-08-06Alignment Is Local: A Paired Diagnostic for GUI Agents under User-Side Persuasion · 2026-07-314.0202026-08-06
64Chunrong FangBreaking Customized LLMs for Coding: Automated Red Teaming for Instruction Backdoor Attacks · 2026-08-06Execution-Grounded Security Testing for Coding Agents in Software Engineering Pipelines · 2026-06-014.0202026-08-06
65David SchmotzStealing Reasoning Traces from Proprietary LLM APIs · 2026-08-10ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D · 2026-07-214.0202026-08-10
66Di WangProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions · 2026-08-11Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs · 2026-08-104.0202026-08-11
67Dingyan ShangSafety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks · 2026-07-30Self-Commitment Latency: A Reward-Free Probe for Prompted Implicit Hacking · 2026-06-044.0202026-07-30
68Dongbin NaWhen Are Reasoning-Based Guardrails Not Efficient? ResponseGuard: A Fast Vision-Language Guard f · 2026-07-23 1stDo Safety Guardrails Need to Reason? LeanGuard: A Fast and Light Approach for Robust Moderation · 2026-06-25 1st4.0202026-07-23
69Feng ChenREDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems · 2026-08-11Stress Testing Concept Erasure with Large Language Model Agents · 2026-07-204.0202026-08-11
70Hangtao ZhangTYPO: Instruction-Dense Visual Jailbreaks against Commercial Closed-Source Image-Generation Mode · 2026-07-27PVDetector: Detecting Prompt Injection Attacks on Purpose-Specific LLM Agents through Policy-Vio · 2026-07-144.0202026-07-27
71Jan DubinskiValue Leakage: An LLM's Answers Are Silently Shaped by Its Own Values · 2026-07-15Negation Neglect: When models fail to learn negations in training · 2026-05-134.0202026-07-15
72Jannik BrinkmannSynthetic Persona Pretraining: Alignment from Token Zero · 2026-08-13Mood Matters: How Syntactic Sensitivity Undermines Safety Alignment · 2026-08-054.0202026-08-13
73Jiawei ChenTrace, Verify, and Correct: A Training-Free Framework for Spatial Reasoning in Multimodal LLMs · 2026-08-05The Verification Horizon: No Silver Bullet for Coding Agent Rewards · 2026-06-244.0202026-08-05
74Jie ZhangBenchmarking Cyberattack Detection in Electric Vehicle Charging Infrastructure with Benign User · 2026-08-11An Early Warning of Emerging Biosecurity Risks in Frontier LLMs · 2026-07-204.0202026-08-11
75Jing LiWhen Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Sk · 2026-08-09ASRU: Activation Steering Meets Reinforcement Unlearning for Multimodal Large Language Models · 2026-05-154.0202026-08-09
76Jinqiu YangChecked-In Secret Detection: Strings Are All You Need · 2026-08-05IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests · 2026-07-224.0202026-08-05
77Joe BentonDiffuse AI Control on Fuzzy Tasks · 2026-06-08Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning · 2026-05-22Removing Sandbagging in LLMs by Training with Weak Supervision · 2026-04-234.0302026-06-08
78Junkai ChenExploring and Bridging Knowledge Holes in Unlearned Multimodal Large Language Models · 2026-08-03Toward Fine-Grained Forgetting:Attribute Unlearning for Multimodal Large Language Models · 2026-08-024.0202026-08-03
79Juntao DaiA Blind Spot in Alignment: Quantifying Biosecurity Risks in Large Language Models · 2026-08-03When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agents · 2026-06-184.0202026-08-03
80Li ZengTYPO: Instruction-Dense Visual Jailbreaks against Commercial Closed-Source Image-Generation Mode · 2026-07-27PVDetector: Detecting Prompt Injection Attacks on Purpose-Specific LLM Agents through Policy-Vio · 2026-07-144.0202026-07-27