Rising2 sources· last seen 3h ago· first seen 2d ago
Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits
Diffusion Large Language Models (DLLMs) replace autoregressive next-token prediction with iterative parallel denoising, yet their internal safety mechanisms remain poorly understood. In this work, we investigate DLLMs both as targets and as adversaries, exposing mechanistic vulnerabilities in diffus
Lead: arXivBigness: 44diffusionllmstargetsadversariesmechanistic
📡 Coverage
50
2 news sources
🟠 Hacker News
18
2 pts, 0 comments
🔴 Reddit
0
📈 Google Trends
0
Full methodology: How scoring works
Receipts (all sources)
Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits
HACKERNEWS · Hacker News · 3h ago · ▲ 2
score 154
Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits
ARXIV · arXiv · 2d ago
score 14
Diffusion Large Language Models (DLLMs) replace autoregressive next-token prediction with iterative parallel denoising, yet their internal safety mechanisms remain poorly understood. In this work, we investigate DLLMs both as targets and as adversaries, exposing mechanistic vulnerabilities in diffus