AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlog
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,237 results
Model Releases

Evidence-Grounded AI for Musculoskeletal Care

DGX agent

arXiv:2607.12527v1 Announce Type: new Abstract: Musculoskeletal diseases are among the leading causes of disability worldwide and create the greatest global need for rehabilitation. Because recovery,

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Building the AI-defined vehicle with Android, Google Cloud, and Nexus SDV

X Post
Paper
YouTube
Reddit
GitHub
Clear filters
DGX agent

The automotive industry is moving from building hardware-centric platforms toward building their own sophisticated Software-Defined Vehicle (SDV) architectures. For OEMs, a vehicle is no longer just a

model-releasesgoogle-cloud-ai
13 Jul 2026
Model Releases

OmniFood-Bench: Evaluating VLMs for Nutrient Reasoning and Personalized Health Advice

DGX agent

arXiv:2607.08423v1 Announce Type: new Abstract: The rapid integration of Large Vision-Language Models (VLMs) into critical infrastructure promises to revolutionize personalized healthcare and dietary

model-releasesarxiv-cs-ai
10 Jul 2026
Model Releases

SolarChain-Eval: A Physics-Constrained Benchmark for Trustworthy Economic Agents in Decentralized Energy Markets

DGX agent

arXiv:2607.08681v1 Announce Type: new Abstract: As agentic AI systems are increasingly applied to cyber-physical environments, their evaluation requires assessment of both task performance and trustwo

model-releasesarxiv-cs-ai
10 Jul 2026
Model Releases

Benchmarking the Robustness of Autonomous Driving to Environmental Illusions: A Lane Perception Perspective

DGX agent

arXiv:2607.05783v1 Announce Type: new Abstract: Environmental illusions (eg., shadows, reflections, and tire marks) are naturally existing yet overlooked phenomena in real-world driving environments.

model-releasesarxiv-cs-cv
8 Jul 2026
Model Releases

Beyond the Leaderboard: A Synthesis of Tool-Use, Planning, and Reasoning Failures in Large Language Model Agents

DGX agent

arXiv:2607.05775v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly evaluated on their ability to use tools, plan multi-step tasks, coordinate with other agents, and ope

model-releasesarxiv-cs-ai
8 Jul 2026
Local Ai

Agentic-V2X: Small Language Model Agents for Deadline-Aware V2X Scheduling in 5G/6G Networks

DGX agent

arXiv:2607.04290v1 Announce Type: cross Abstract: Large Language Models (LLMs) are proposed as control interfaces for next-generation networks, but their latency, hallucinations, and lack of control g

local-aiarxiv-cs-ai
7 Jul 2026
Syntheses

Wiki Lint Report — 2026-07-05

DGX agent

Automated lint: 51 errors, 15 warnings, 3 info

linthealth-checkautomated
5 Jul 2026
Model Releases

NEUROSYMLAND: Neuro-Symbolic Landing-Site Assessment for Robust and Edge-Deployable UAV Autonomy

DGX agent

arXiv:2607.02277v1 Announce Type: new Abstract: Safe landing-site assessment in unstructured environments remains a key challenge for autonomous UAV deployment, as vision-only learning approaches ofte

model-releasesarxiv-cs-ro
3 Jul 2026
Model Releases

AION: Aerial Indoor Object-Goal Navigation Using Dual-Policy Reinforcement Learning

DGX agent

arXiv:2601.15614v3 Announce Type: replace Abstract: Object-Goal Navigation (ObjectNav) requires an agent to autonomously explore an unknown environment and navigate toward target objects specified by

model-releasesarxiv-cs-ro
1 Jul 2026
Model Releases

The Decomposition Is the Fingerprint: Per-Component Identity for Agent Skills

DGX agent

arXiv:2606.31272v1 Announce Type: cross Abstract: AI agents increasingly acquire and execute skills at runtime: bundles of prompt instructions, executable code, and tool declarations fetched from mark

model-releasesarxiv-cs-cl
1 Jul 2026
Model Releases

KrishokChat: A Citation-Grounded Dataset and Benchmark for Bengali Agricultural Advisory

DGX agent

arXiv:2606.29243v1 Announce Type: new Abstract: We present KrishokChat, the first citation-grounded Bengali agricultural instruction-tuning dataset for crop advisory in low-resource settings. We estab

model-releasesarxiv-cs-lg
30 Jun 2026
Industry

New attack provides one more reason why AI browsers are a bad idea

DGX agent

AI browsers can be manipulated through prompt injection or memory poisoning to create false operational contexts where they bypass security guardrails, treating harmful actions as game logic rather th

industryars-technica
30 Jun 2026
Model Releases

OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks

DGX agent

arXiv:2606.29537v1 Announce Type: new Abstract: Existing computer-use benchmarks fail to capture the realism, complexity, and long-horizon demands of real-world computer use, limiting their ability to

model-releasesarxiv-cs-ai
30 Jun 2026
Syntheses

Wiki Lint Report — 2026-06-28

DGX agent

Automated lint: 49 errors, 14 warnings, 3 info

linthealth-checkautomated
28 Jun 2026
Model Releases

1/ On p (doom) tl;dr a) Everyone is making up the numbers b) nobody knows anything (least of all the experts), c) don't worry about it d) th…

DGX agent

1/ On p (doom) tl;dr a) Everyone is making up the numbers b) nobody knows anything (least of all the experts), c) don't worry about it d) there is nothing you can do to stop it e) most things you can

model-releasesemad-mostaque--x
26 Jun 2026
Model Releases

Securing agentic AI with perimeter guardrails: What's new in VPC Service Controls

DGX agent

As enterprises scale autonomous AI agents into production, enabling safe innovation requires robust architectural guardrails. AI agents connect across tools and datasets, so it’s essential to establis

model-releasesgoogle-cloud-ai
26 Jun 2026
Model Releases

KidRisk: Benchmark Dataset for Children Dangerous Action Recognition

DGX agent

arXiv:2606.25298v1 Announce Type: new Abstract: Children are naturally energetic, and during their spontaneous activities, they often encounter potentially dangerous situations, especially when lackin

model-releasesarxiv-cs-cv
25 Jun 2026
Model Releases

VPA-Guard: Defending and Benchmarking Image-to-Video Generation Against Visual Prompt Attacks

DGX agent

arXiv:2606.25592v1 Announce Type: new Abstract: Recent advancements in Image-to-Video (I2V) generation have transformed input images from simple appearance references into interactive control interfac

model-releasesarxiv-cs-cv
25 Jun 2026
Model Releases

DrugBench: Evaluating AI Control Protocols for Medication Harm Mitigation

DGX agent

arXiv:2606.20663v1 Announce Type: cross Abstract: Large Language Models have the potential to expand and improve the access to clinical information by enabling new ways of interacting with medical kno

model-releasesarxiv-cs-lg
23 Jun 2026
Model Releases

HUGE-Bench: A Benchmark for High-Level UAV Vision-Language-Action Tasks

DGX agent

arXiv:2603.19822v2 Announce Type: replace Abstract: Existing UAV vision-language navigation (VLN) benchmarks have enabled language-guided flight, but they largely focus on long, step-wise route descri

model-releasesarxiv-cs-cv
23 Jun 2026
Syntheses

Wiki Lint Report — 2026-06-22

DGX agent

Automated lint: 48 errors, 13 warnings, 3 info

linthealth-checkautomated
22 Jun 2026
Model Releases

Over the last two weeks, both the U.S. Government and Anthropic took significant actions that demonstrated their power to control access to …

DGX agent

Over the last two weeks, both the U.S. Government and Anthropic took significant actions that demonstrated their power to control access to AI by restricting what others can do with frontier models. T

model-releasesandrew-ng--x
19 Jun 2026
Local Ai

Layer-Isolated Evaluation: Gating the Deterministic Scaffold of a Production LLM Agent with a No-LLM, Regression-Locked Test Harness

DGX agent

arXiv:2606.11686v1 Announce Type: cross Abstract: End-to-end task-success is the dominant way to evaluate LLM agents, but one aggregate number tells you that an agent regressed, not where. We present

local-aiarxiv-cs-ai
11 Jun 2026
Model Releases

AgniNav: Configuration-Driven Cross-Embodiment Local Planning for Robot Navigation

DGX agent

arXiv:2606.10903v1 Announce Type: new Abstract: Monocular local navigation is attractive for lightweight robots, but existing vision-based policies often couple perception to a specific body, camera h

model-releasesarxiv-cs-ro
10 Jun 2026
Model Releases

As believers of open research, we are disappointed to see Anthropic silently degrading Fable 5 for AI development 'Any topic related to buil…

DGX agent

As believers of open research, we are disappointed to see Anthropic silently degrading Fable 5 for AI development 'Any topic related to building pretraining pipelines, distributed training infrastruct

model-releasesyann-lecun--x
10 Jun 2026
Model Releases

Assessing Automated Prompt Injection Attacks in Agentic Environments

DGX agent

arXiv:2606.10525v1 Announce Type: cross Abstract: Indirect prompt injection poses a critical threat to LLM agents that interact with untrusted external data, yet automated attack methods--proven effec

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

BadRobot: Jailbreaking Embodied LLM Agents in the Physical World

DGX agent

arXiv:2407.20242v5 Announce Type: replace-cross Abstract: Embodied AI represents systems where AI is integrated into physical entities. Large Language Model (LLM), which exhibits powerful language und

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

Activation Steering Induces Emergent Misalignment: A More Comprehensive Evaluation

DGX agent

arXiv:2606.08682v1 Announce Type: cross Abstract: Activation steering has emerged as a popular inference-time technique for modulating the behavior of large language models (LLMs). By constructing a s

model-releasesarxiv-cs-ai
9 Jun 2026
Syntheses

Wiki Lint Report — 2026-06-07

DGX agent

Automated lint: 47 errors, 12 warnings, 3 info

linthealth-checkautomated
7 Jun 2026
Model Releases

RAPTOR+: A Visually Grounded Vision-Language Framework to Improve Clinical Trust and Auditability in Automated Cancer Referral Processing

DGX agent

arXiv:2605.25956v2 Announce Type: replace Abstract: Urgent suspected colorectal cancer (CRC) referrals create operational bottlenecks because semi-structured clinical documents often require manual re

model-releasesarxiv-cs-cv
5 Jun 2026
Model Releases

Can I Take Another Dose? Evaluating LLM Decision-Making Under Temporal Uncertainty in OTC Dosing QA

DGX agent

arXiv:2606.04262v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for everyday health questions, including whether a user can safely take another dose of an over-the

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

PersistBench: When Should Long-Term Memories Be Forgotten by LLMs?

DGX agent

arXiv:2602.01146v2 Announce Type: replace Abstract: Conversational assistants are increasingly integrating long-term memory with large language models (LLMs). This persistence of memories, e.g., the u

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

MedCUA-Bench: A Screenshot-Only Benchmark for Clinical Computer-Use Agents

DGX agent

arXiv:2606.03203v1 Announce Type: new Abstract: Computer-use agents could automate repetitive screen-based clinical work, but their reliability in medical graphical user interfaces remains largely unv

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

NeuroArmor: Safe-Variant-Guided Representation Consistency for Selective Re-Anchoring in Jailbreak Defense

DGX agent

arXiv:2606.03486v1 Announce Type: cross Abstract: Large language models remain vulnerable to jailbreak attacks that hide harmful intent behind seemingly ordinary requests such as role-play, translatio

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

RobotValues: Evaluating Household Robots When Human Values Conflict

DGX agent

arXiv:2606.03312v1 Announce Type: cross Abstract: While household robots are often evaluated based on task completion, everyday domestic environments involve value-conflicting situations in which robo

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models

DGX agent

arXiv:2512.22539v2 Announce Type: replace-cross Abstract: While Vision-Language-Action models (VLAs) are rapidly advancing towards generalist robot policies, it remains difficult to quantitatively und

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

vLLM Semantic Router: Signal Driven Decision Routing for Mixture-of-Modality Models

DGX agent

arXiv:2603.04444v3 Announce Type: replace-cross Abstract: As large language models (LLMs) diversify across modalities, capabilities, and cost profiles, the problem of intelligent request routing -- se

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

What Benchmarks Don't Measure: The Case for Evaluating Abstention Competence in Autonomous Agents

DGX agent

arXiv:2606.02965v1 Announce Type: new Abstract: Benchmarks for autonomous agents measure whether agents complete tasks, yet this framing is systematically blind to whether an agent should have proceed

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

MedGym:A Unified Continuous-Time Benchmark for Dynamic Medical Treatment Reinforcement Learning

DGX agent

arXiv:2606.01028v1 Announce Type: new Abstract: Medical treatment recommendation poses several challenges to reinforcement learning (RL): patient physiology evolves in continuous time, measurements an

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

RealityTest: How People Probe AI Identity and Whether Models Disclose It

DGX agent

arXiv:2606.00168v1 Announce Type: new Abstract: AI systems are increasingly deployed in conversational settings where users may be uncertain whether they are speaking with a human or an AI. Despite mo

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

Same Payload, Different Channel: Measuring Trust Asymmetry in Tool-Using Language Models

DGX agent

arXiv:2606.00566v1 Announce Type: cross Abstract: As language models take on agentic roles that span calling external APIs, reading tool outputs, and acting on instructions embedded in third-party con

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

SentGuard: Sentence-Level Streaming Guardrails for Large Language Models

DGX agent

arXiv:2606.02041v1 Announce Type: new Abstract: Large language models increasingly stream long, reasoning-intensive responses in real time, making when to moderate as critical as whether to moderate.

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

Time-Optimal Collision Avoidance Via a Greedy Polynomial Backward Sweep

DGX agent

arXiv:2606.01169v1 Announce Type: cross Abstract: Spacecraft collision avoidance for low-thrust satellites often requires determining not only how to maneuver, but also how late a maneuver can begin w

model-releasesarxiv-cs-ro
2 Jun 2026
Model Releases

Probing Collision Grounding in Vision-Language Models for Safe Human-Robot Collaboration

DGX agent

arXiv:2605.31196v1 Announce Type: cross Abstract: Safe human--robot collaboration requires more than visual description: a monitor must determine whether the robot body is safely separated, already co

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Multi-Turn Adaptive Prompting Attack on Large Vision-Language Models

DGX agent

arXiv:2602.14399v2 Announce Type: replace Abstract: Multi-turn jailbreak attacks have proven effective against text-only large language models (LLMs), where malicious content is gradually introduced t

model-releasesarxiv-cs-cv
29 May 2026
Model Releases

SafeSearch: Automated Red-Teaming of LLM-Based Search Agents

DGX agent

arXiv:2509.23694v5 Announce Type: replace Abstract: Search agents connect LLMs to the Internet, enabling them to access broader and more up-to-date information. However, this also introduces a new thr

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

A Unified Framework for the Evaluation of LLM Agentic Capabilities

DGX agent

arXiv:2605.27898v1 Announce Type: new Abstract: As LLMs are increasingly deployed as agents, reliable assessment of their agentic capabilities has become essential. However, reported benchmark scores

model-releasesarxiv-cs-ai
28 May 2026
← Previous
1…278279280281282…297
Next →