AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlog
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,351 results
Safety

bioMoR: Biology-Guided Mixture-of-Recursions for Effective Genomic Learning

DGX agent

arXiv:2608.06727v1 Announce Type: new Abstract: Transformer models for high-dimensional omics analysis process thousands of genes or pathways, although only a subset requires deep computation. Mixture

safetyarxiv-cs-ai
10 Aug 2026
Safety
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

From Cheap Fakes to Pure Synthesis: Addressing the New Era of T2V Fake News Videos

DGX agent

arXiv:2608.06732v1 Announce Type: new Abstract: Recent text-to-video (T2V) generation models enable fake news videos to be synthesized from scratch, shifting the threat beyond cheap fakes assembled fr

safetyarxiv-cs-ai
10 Aug 2026
Safety

SoRoMoX: Fast, Differentiable, and Parallelizable Soft Robot Models

DGX agent

arXiv:2608.06650v1 Announce Type: cross Abstract: Reduced-order models based on Cosserat-rod theory are now well established, and modeling theory is no longer the primary bottleneck in soft-robot cont

safetyarxiv-cs-ai
10 Aug 2026
Safety

A System for Train Condition Monitoring and Structural Health Assessment of Rail Vehicles

DGX agent

arXiv:2608.05221v1 Announce Type: new Abstract: The ongoing digitalization of rail systems and the increasing use of artificial intelligence (AI) are fundamentally transforming the design, operation,

safetyarxiv-cs-ro
7 Aug 2026
Safety

DreamGuard: Efficient Runtime Guardrail for LLM Agents via Risk-Aware World Model

DGX agent

arXiv:2608.05695v1 Announce Type: new Abstract: As large language model (LLM) agents increasingly invoke external tools and interact with real-world systems, unsafe actions may cause irreversible cons

safetyarxiv-cs-ai
7 Aug 2026
Safety

Looking in the Mirror: Introspecting Side-Effect Misalignments Induced by Fine-Tuning

DGX agent

arXiv:2608.04347v1 Announce Type: new Abstract: Fine-tuning enables a source model to acquire desired capabilities and behaviors in a target domain while retaining much of its general-purpose competen

safetyarxiv-cs-lg
6 Aug 2026
Safety

A Blind Spot in Alignment: Quantifying Biosecurity Risks in Large Language Models

DGX agent

arXiv:2608.02684v1 Announce Type: cross Abstract: Large Language Models (LLMs) are accelerating biological research, yet this same capability poses a critical biosecurity threat: models that assist in

safetyarxiv-cs-ai
5 Aug 2026
Safety

LatentGuard: Efficient and Inspectable Latent Reasoning for LLM Safeguards

DGX agent

arXiv:2608.03838v1 Announce Type: new Abstract: Reasoning-based guard models improve LLM safeguards, but decoding explicit rationales for every interaction makes them costly to deploy. Although latent

safetyarxiv-cs-ai
5 Aug 2026
Safety

Poisoning Prompt-Guided Sampling in Video Large Language Models

DGX agent

arXiv:2509.20851v2 Announce Type: replace Abstract: Video Large Language Models (VideoLLMs) are increasingly deployed as automated moderators on user-generated video platforms, where a few unwatched s

safetyarxiv-cs-cv
5 Aug 2026
Safety

PULSE: An Executable Contract Language for Spatiotemporal Knowledge Graph Engineering

DGX agent

arXiv:2608.02630v1 Announce Type: new Abstract: Knowledge graph engineering often distributes accepted state, observations, constraints, processes, and hypothetical scenarios across artifacts whose co

safetyarxiv-cs-ai
5 Aug 2026
Safety

HapticVLA: Contact-Rich Manipulation via Vision-Language-Action Model without Inference-Time Tactile Sensing

DGX agent

arXiv:2603.15257v2 Announce Type: replace Abstract: Tactile sensing is a crucial capability for Vision-Language-Action (VLA) architectures, as it enables dexterous and safe manipulation in contact-ric

safetyarxiv-cs-ro
4 Aug 2026
Model Releases

MedPRESS: A Multi-turn Benchmark for Patient-Pressure-Induced Medical Sycophancy in LLMs

DGX agent

arXiv:2608.02520v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for health-related advice. Existing research measures their safety with static questions rather than

model-releasesarxiv-cs-cl
4 Aug 2026
Model Releases

Quality-Diversity Red-Teaming: Automated Generation of High-Quality and Diverse Attackers for Large Language Models

DGX agent

arXiv:2506.07121v2 Announce Type: replace Abstract: Ensuring the safety and robustness of large language models (LLMs) is a fundamental challenge and a critical prerequisite for the responsible deploy

model-releasesarxiv-cs-lg
4 Aug 2026
Safety

Self-Improving Large Language Models via Progressive Experience Evolution

DGX agent

arXiv:2608.02139v1 Announce Type: new Abstract: Large language models (LLMs) capable of self-improvement require not only effective policy optimization, but also a principled mechanism for transformin

safetyarxiv-cs-cl
4 Aug 2026
Safety

When May a Model Replace the Experiment? Audits, Licenses, and the Price of Trust in Surrogate-Driven Design

DGX agent

arXiv:2608.01378v1 Announce Type: new Abstract: Design campaigns in chemistry, materials science, and machine learning share a bottleneck: determining how good a candidate truly is requires an expensi

safetyarxiv-cs-lg
4 Aug 2026
Safety

CAGE: Certified Authorization under Typed-Return Uncertainty for Tool-Using Agents

DGX agent

arXiv:2607.29190v1 Announce Type: new Abstract: Tool-using LLM agents act on typed tool returns, records pairing provenance and categorical fields with numerical values. Runtime permission gates gener

safetyarxiv-cs-ai
3 Aug 2026
Safety

Active Lubrication of Transluminal Medical Instruments

DGX agent

arXiv:2506.07225v2 Announce Type: replace-cross Abstract: Transluminal minimally invasive surgery uses natural orifices and small incisions to access internal anatomical structures, promoting quicker

safetyarxiv-cs-ro
31 Jul 2026
Safety

Ask don't tell: Reducing sycophancy in large language models

DGX agent

arXiv:2602.23971v4 Announce Type: replace-cross Abstract: Sycophancy, the tendency of large language models to favour user-affirming responses over critical engagement, has been identified as an align

safetyarxiv-cs-ai
31 Jul 2026
Safety

Corrigible Assistance in One Round: Pragmatic-Pedagogic Best Response

DGX agent

arXiv:2607.27508v1 Announce Type: new Abstract: Assistance games formalize human-robot collaboration under asymmetric information: the human knows the goal, while the robot must infer it from observat

safetyarxiv-cs-ro
31 Jul 2026
Safety

EmoFeedback^2: Reinforcement of Continuous Emotional Image Generation via LVLM-based Reward and Textual Feedback

DGX agent

arXiv:2511.19982v3 Announce Type: replace Abstract: Continuous emotional image content generation (C-EICG) is emerging rapidly due to its ability to produce images aligned with both user descriptions

safetyarxiv-cs-cv
31 Jul 2026
Safety

FleetScape: A Mixed Reality Sandtable for Spatial Supervision and Control of Scalable Drone Fleets

DGX agent

arXiv:2607.26423v1 Announce Type: cross Abstract: As autonomous drone deployments scale from individual units to coordinated swarms, the human operator's role shifts from direct piloting to high-level

safetyarxiv-cs-ro
30 Jul 2026
Model Releases

Prosody-driven Jailbreaks in Audio LLMs: A Controlled Study and Mechanistic Analysis

DGX agent

arXiv:2607.26541v1 Announce Type: cross Abstract: Audio-capable foundation models enable end-to-end spoken interaction, but they also introduce safety risks beyond transcript content. It remains uncle

model-releasesarxiv-cs-cl
30 Jul 2026
Safety

Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation

DGX agent

arXiv:2607.25489v1 Announce Type: new Abstract: Large language models and multimodal foundation models are enabling medical artificial intelligence (AI) systems to move beyond isolated prediction and

safetyarxiv-cs-cv
29 Jul 2026
Model Releases

Inspect India Evals: An Open Benchmarking Framework for Evaluating Large Language Models in the Indian Linguistic and Cultural Context

DGX agent

arXiv:2607.25375v1 Announce Type: new Abstract: India is a vast nation of over 1.4 billion people, varied by hundreds of diverse and locally specific traditions and cultures and 22 officially recogniz

model-releasesarxiv-cs-cl
29 Jul 2026
Model Releases

Shieldstral

DGX agent

arXiv:2607.25857v1 Announce Type: new Abstract: We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7imes its size on text s

model-releasesarxiv-cs-cl
29 Jul 2026
Safety

SONG: A Photorealistic 3D Gaussian Simulation Platform for Benchmarking Social Navigation

DGX agent

arXiv:2607.25219v1 Announce Type: new Abstract: Social navigation has progressed from simplified 2D environments toward a more general vision-based setting, in which a robot needs to achieve socially

safetyarxiv-cs-ro
29 Jul 2026
Safety

A Computational Ethical Framework for Financial Digital Phenotyping for Mental Health

DGX agent

arXiv:2607.24275v1 Announce Type: cross Abstract: Ethical governance of AI-driven systems is often expressed through high-level principles and static documentation, creating a gap between regulatory r

safetyarxiv-cs-ai
28 Jul 2026
Safety

Adapter Merging Reactivates Latent Reasoning Traces: A Mechanism Analysis

DGX agent

arXiv:2601.18350v5 Announce Type: replace-cross Abstract: Large language models fine-tuned via a two-stage pipeline (domain adaptation followed by instruction alignment) can exhibit non-trivial interf

safetyarxiv-cs-ai
28 Jul 2026
Safety

ARdena: Scenario-driven control of real-time LLM agents

DGX agent

arXiv:2607.22651v1 Announce Type: new Abstract: Large language models (LLMs) have enabled increasingly capable conversational agents, but reliably controlling their behavior in real-time interactive e

safetyarxiv-cs-ai
28 Jul 2026
Safety

Frustratingly Simple Black-Box Adaptation of Language Models via Logit Bias

DGX agent

arXiv:2607.22837v1 Announce Type: cross Abstract: Many organizations aim to adapt language models for internal use, both to improve performance on domain-specific tasks and to address privacy concerns

safetyarxiv-cs-ai
28 Jul 2026
Safety

How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift

DGX agent

arXiv:2607.22676v1 Announce Type: new Abstract: Post-training is a key mechanism for adapting large language models to downstream tasks. While prior work suggests that task adaptation can alter a mode

safetyarxiv-cs-ai
28 Jul 2026
Safety

IJCB-AFMFR 2026: Competition on Adapting Foundation Models for Face Recognition Using Synthetic Training Data

DGX agent

arXiv:2607.24422v1 Announce Type: new Abstract: This paper presents a summary of the Competition on Adapting Foundation Models for Face Recognition Using Synthetic Training Data (AFMFR), held at the 2

safetyarxiv-cs-cv
28 Jul 2026
Safety

Position: AI/ML Deepfake Research is Misaligned with AI-Generated Non-Consensual Intimate Imagery (AIG-NCII)

DGX agent

arXiv:2607.18263v2 Announce Type: replace Abstract: AI-generated non-consensual intimate imagery (AIG-NCII) is not adequately addressed in AI/ML literature regarding AI-generated media, commonly refer

safetyarxiv-cs-ai
28 Jul 2026
Safety

RM-Distiller: Exploiting Generative LLM for Reward Model Distillation

DGX agent

arXiv:2601.14032v2 Announce Type: replace Abstract: Reward models (RMs) play a pivotal role in aligning large language models (LLMs) with human preferences. Due to the difficulty of obtaining high-qua

safetyarxiv-cs-cl
28 Jul 2026
Safety

Self-Authored Verification Is Unreliable in Heuristic Self-Improving Agents

DGX agent

arXiv:2607.24300v1 Announce Type: new Abstract: Self-improving agents accumulate capability by repeatedly rewriting procedural policies, controllers, or heuristic rules. They typically rely on self-au

safetyarxiv-cs-cl
28 Jul 2026
Safety

STAIF: A Stage-wise Optimization for Complex Instruction Following

DGX agent

arXiv:2607.22649v1 Announce Type: new Abstract: Following complex instructions with multiple explicit constraints remains a fundamental challenge for large language models (LLMs). Existing alignment m

safetyarxiv-cs-ai
28 Jul 2026
Safety

Tailored untruths: How personalisation challenges LLM safeguards

DGX agent

arXiv:2510.12993v3 Announce Type: replace Abstract: Large Language Models (LLMs) can generate highly persuasive disinformation, yet little is known about how effectively they personalise it across lan

safetyarxiv-cs-cl
28 Jul 2026
Safety

Unequal Trips, Unequal Places: Diagnosing and Mitigating Delay Inequity in Autonomous Vehicle Fleet Coordination

DGX agent

arXiv:2607.24336v1 Announce Type: new Abstract: City-scale autonomous vehicle fleet coordinators are typically optimized for aggregate travel time, yet fleet averages conceal how delay is distributed

safetyarxiv-cs-ai
28 Jul 2026
Safety

White-hat hacking IS the defense to black-hat hacking. The techniques are the same. How does Dario expect companies to do it if their models refuse?

DGX agent

You patch security holes by intentionally finding them. If the models refuse to do it, how can companies protect themselves against rogue AIs, whether they are Chinese or OpenAI/Anthropic themselves?

safetyr-localllama
28 Jul 2026
Safety

Why Anthropic's battle is meant to poison the wells of open weight models, in 3 steps.

DGX agent

It doesn't solve any problems. Just a few paragraphs above, he says he fears that authoritarian states (he names China, and possibly others) can use their models to do evil stuff. And surely enough, m

safetyr-localllama
28 Jul 2026
Safety

We recognize there are a lot of questions and speculative details circulating related to the Hugging Face incident. This is an unprecedented…

DGX agent

We recognize there are a lot of questions and speculative details circulating related to the Hugging Face incident. This is an unprecedented incident, and we think it marks an important moment for AI

safetyopenai--x
25 Jul 2026
Safety

Interaction Dynamics Modeling and Predictive Control for Safe Steerable Catheter--Tissue Interaction

DGX agent

arXiv:2607.20939v1 Announce Type: cross Abstract: Safe steerable catheter control is fundamentally a problem of interaction dynamics: the tip must follow a planned motion, remain compliant against mov

safetyarxiv-cs-ai
24 Jul 2026
Safety

OpenAI and Anthropic are now competing to see who can make the scariest AI for marketing purposes. What could possibly go wrong?

DGX agent

Pedro Domingos noted that OpenAI and Anthropic are competing to develop the scariest AI for marketing, raising concerns about the potential misuse of persuasive machine‑learning models. His comment hi

safetygary-marcus--x
24 Jul 2026
Safety

Robust Critics: Defending LLMs Against Multi-Turn Attacks

DGX agent

arXiv:2607.20472v1 Announce Type: new Abstract: When a user asks a language model something harmful, is it a genuine attack or a misunderstood but well-meaning question? This ambiguity is one of the c

safetyarxiv-cs-ai
24 Jul 2026
Safety

Safe and Scalable Multi-Drone Payload Transport via CBF-based Reinforcement Learning with Zero-Shot Sim-to-Real Transfer

DGX agent

arXiv:2607.20665v1 Announce Type: new Abstract: Multi-drone payload transportation has emerged as a promising research paradigm with potential applications in construction, logistics, and disaster res

safetyarxiv-cs-ro
24 Jul 2026
Safety

Formal Foundations for Known Good Reliable Die Screening in Chiplet-Based AI Systems-on-Chip

DGX agent

arXiv:2607.20141v1 Announce Type: cross Abstract: The rapid growth of chiplet-based artificial intelligence systems-on-chip (SoCs) has exposed a fundamental gap in semiconductor test methodology. Exis

safetyarxiv-cs-ai
23 Jul 2026
Safety

No Training, Better Flights: Test-Time Scaled VLMs for UAV Navigation

DGX agent

arXiv:2607.19288v1 Announce Type: new Abstract: Test-time scaling offers a promising method to improve the inference performance of Vision-Language Models (VLMs) without additional training. Existing

safetyarxiv-cs-cv
23 Jul 2026
Safety

OPIUM: Mitigating Steering Externalities and Over-Refusal via Dual Objective Latent Optimization

DGX agent

arXiv:2607.19806v1 Announce Type: cross Abstract: Activation steering provides a lightweight mechanism for controlling large language models at inference time, but steering vectors can have unintended

safetyarxiv-cs-ai
23 Jul 2026
← Previous
1…3334353637…299
Next →