AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,237 results
16 Jul 2026

HRIBench: Benchmarking Interaction-Centric Human-Robot Collaboration

Model ReleasesDGX agent

arXiv:2607.13056v1 Announce Type: cross Abstract: Current vision-language-action (VLA) benchmarks primarily evaluate isolated manipulation skills while leaving human-robot interaction structure largel

MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation

Model ReleasesDGX agent

arXiv:2607.09142v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in online medical consultation, yet existing benchmarks remain poorly aligned with real clini

Not All Needles Are Found: How Fact Distribution and Don't Make It Up Prompts Shape Retrieval, Reasoning, and Hallucination in Long-Context LLMs

Model Releases
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2601.02023v2 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) increasingly utilize massive context windows as working memory for autonomous tasks, their reliability fluctua

nuTruck: Benchmarking Autonomous Driving Planning for Distributed Electric-drive Trucks

Model ReleasesDGX agent

arXiv:2607.13704v1 Announce Type: new Abstract: The dominance of traditional rule-based methods in autonomous driving has gradually been replaced by learning-based approaches. While learning-based pla

Securing LLMs in the Wild: Privacy and Security Challenges at the Edge

Model ReleasesDGX agent

arXiv:2607.13088v1 Announce Type: cross Abstract: Large Language Models (LLMs) are rapidly moving from research settings into the wild, deployed on enterprise infrastructure, personal devices, and edg

15 Jul 2026

Agents Don't Just Agree, They Remember: Benchmarking Persistent Sycophancy in Stateful Personal Agents

Model ReleasesDGX agent

arXiv:2607.10526v2 Announce Type: replace Abstract: Stateful personal agents increasingly maintain long-term user profiles, episodic memories, and reusable skills. This persistence turns conversationa

Are we Merging the Right Models? Impact of Expert Training Duration on Model Merging for LLMs

Model ReleasesDGX agent

arXiv:2607.11997v1 Announce Type: cross Abstract: Multi-task model merging combines separately trained expert models into a single model that handles all tasks without co-training. Standard practice m

Breaking Deja Vu: Independent Auditing of Visual Place Recognition through Vision-Language Reasoning

Model ReleasesDGX agent

arXiv:2607.12818v1 Announce Type: new Abstract: Visual place recognition (VPR) is a key enabler of accurate localization and long-term autonomous navigation in robotics applications, such as loop clos

Build a Multi-Camera 3D Tracking Application with NVIDIA DeepStream 9.1 Skills

HardwareDGX agent

NVIDIA DeepStream 9.1 introduces AutoMagicCalib (AMC) and Multi‑View 3D Tracking (MV3DT) to automate camera calibration and maintain consistent 3‑D object IDs across multiple calibrated cameras, reduc

Continual Learning with Elastic Regularization and Synthetic Replay for Federated MLLM Fine-Tuning

Model ReleasesDGX agent

arXiv:2607.12112v1 Announce Type: cross Abstract: Federated fine-tuning of Multimodal Large Language Models (MLLMs) across distributed networks enables privacy-sensitive adaptation to evolving data st

Epistemic Stance Flexibility Probing: Measuring Prompt-Conditioned Register Shift in Large Language Models

Model ReleasesDGX agent

arXiv:2607.12739v1 Announce Type: new Abstract: A language model may be asked either what experts believe about a contested claim or what it believes about the claim itself. A trustworthy conversation

How Inference Compute Shapes Frontier LLM Evaluation

Model ReleasesDGX agent

arXiv:2606.17930v2 Announce Type: replace Abstract: AI evaluations are shifting toward harder tasks that benefit from longer trajectories involving tool use and iterative problem solving. As a result,

How Many Tasks Are Enough for Agent Benchmark Decisions? A Replay Analysis of Public LLM Agent Benchmarks

Model ReleasesDGX agent

arXiv:2607.12338v1 Announce Type: new Abstract: Agent benchmarks often compare two agents after all tasks have run, but costly evaluations make partial runs tempting. A task fraction alone does not sh

How to Realize Recursively Self-Improving Agents and Personal Singularity: A Goal-, Scope-, Tool-, and Benchmark-Driven Multi-Agent Architecture

Model ReleasesDGX agent

arXiv:2607.12254v1 Announce Type: new Abstract: Large language model (LLM) agents can increasingly plan, use tools, maintain memory, and execute long-horizon tasks. These advances motivate two linked

I built a new attention mechanism (wave field) — runs 128K context where standard attention OOMs, 80+ tok/s on laptop CPU

Model ReleasesDGX agent

Hey r/LocalLLaMA — solo researcher here. I built a new attention architecture and want independent testers. Wave Field LLM replaces O(N²) dot-product attention with FFT wave convolution on a field. Tr

Knowledge- and Gradient-Guided Reinforcement Learning for Parametrized Action Markov Decision Processes

Model ReleasesDGX agent

arXiv:2607.12924v1 Announce Type: new Abstract: In this paper, we study Reinforcement Learning in Parametrized Action Markov Decision Processes (PAMDP), where each decision consists of a symbolic acti

Wiki Lint Report — 2026-07-15

SynthesesDGX agent

Automated lint: 26 errors, 6728 warnings, 3 info

Ontology-Amplified Distillation and Contextuality Auditing for Sovereign Enterprise Language Models: A Combined Proof-of-Mechanism and Negative-Results Method Study

Model ReleasesDGX agent

arXiv:2607.11948v1 Announce Type: new Abstract: Regulated financial institutions operating under data-residency rules need tenant-owned language models that can run inside the institution's perimeter.

🆕This Year In Claude https://www.youtube.com/watch?v=uU5Gv2h8-9g @simonw chats with @_catwu and @trq212 about the state of: - @claudeai Cod…

Model ReleasesDGX agent

🆕This Year In Claude https://www.youtube.com/watch?v=uU5Gv2h8-9g @simonw chats with @_catwu and @trq212 about the state of: - @claudeai Code - Claude Fable - @anthropicai culture & product strategy -

Who touches every token that flows through @anthropic? It’s not any one model, but it’s @katelyn_lesse, @angjiang and the platform team. A y…

Model ReleasesDGX agent

Who touches every token that flows through @anthropic? It’s not any one model, but it’s @katelyn_lesse, @angjiang and the platform team. A year ago, it was just a messages API. Today, their platform s

13 Jul 2026

Empowering India’s next generation of innovators with ATL Saathi

Model ReleasesDGX agent

ATL Saathi is an initiative aimed at boosting scientific innovation and learning in India through AI-powered tools and programs. Launched in February 2026, the program focuses on accelerating research

🆕 In Code They Act, In Proof We Trust — Erik Meijer last year, @solomonstre defined agents as 'an LLM that's wrecking its environment in a …

Model ReleasesDGX agent

🆕 In Code They Act, In Proof We Trust — Erik Meijer last year, @solomonstre defined agents as 'an LLM that's wrecking its environment in a loop', and @simonw coined the Lethal Trifecta for agents, tha

Key findings from the 2026 Public Sector M-Trends report and beyond

Model ReleasesDGX agent

In 2026, the public sector is no longer defending a traditional perimeter. Instead, they are defending a complex web of interconnected trust relationships against adversaries that now operate at machi

10 Jul 2026

OmniOPSD: Rationale-Privileged On-Policy Self-Distillation for Affective Computing

Local AiDGX agent

arXiv:2606.15920v2 Announce Type: replace Abstract: Reinforcement learning for multimodal large language models (MLLMs) is often hindered by severe reward sparsity in complex reasoning tasks. This cha

Secure Decentralized Federated Learning via Gossip and Virtual Voting

Model ReleasesDGX agent

arXiv:2607.08651v1 Announce Type: new Abstract: Decentralized federated learning (DFL) removes the central server by letting nodes exchange model updates through peer-to-peer gossip, but existing goss

Shift & Drift: A Zero-Shot Benchmark for Generalizable and Robust Autonomous Driving Motion Planning

Model ReleasesDGX agent

arXiv:2607.07844v1 Announce Type: cross Abstract: While closed-loop motion planners trained on large-scale, object-level datasets, e.g., nuPlan, demonstrate strong in-distribution (ID) performance, th

9 Jul 2026

Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI Agents

Model ReleasesDGX agent

arXiv:2607.07474v1 Announce Type: cross Abstract: Agentic red-teaming benchmarks report whether an injected agent was compromised as a single bit: the attack succeeded, or it did not. We argue that th

Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages

Model ReleasesDGX agent

arXiv:2607.06596v1 Announce Type: cross Abstract: Trusted monitoring is a central defense in AI control: a cheaper trusted model scores an untrusted model's actions for sabotage, and the most suspicio

Multi-Agent Robotic Control with Onboard Vision-Language Models

Model ReleasesDGX agent

arXiv:2607.07403v1 Announce Type: cross Abstract: Vision Language Models (VLMs) and Vision Language Action (VLA) models have shown promise in robotic control. Yet, they face significant challenges reg

SmartHomeSecure: Automated Detection and Repair of Smart Home Configuration Errors Using Large Language Models

Model ReleasesDGX agent

arXiv:2607.06748v1 Announce Type: cross Abstract: Smart home automation platforms increasingly rely on user-authored YAML configuration files to define device behaviors, but these files are prone to s

8 Jul 2026

aiAuthZ: Off-Host, Identity-Bound Authorization for AI Agents

Model ReleasesDGX agent

arXiv:2607.05518v1 Announce Type: cross Abstract: AI agents issue tool calls on the basis of text they cannot verify, so any party who controls part of the context can forge the appearance of authorit

Great writeup from the University of Oxford. It's a taxonomy of LLM-based agent limitations. Good read for anyone shipping with agents. Benc…

Model ReleasesDGX agent

Great writeup from the University of Oxford. It's a taxonomy of LLM-based agent limitations. Good read for anyone shipping with agents. Benchmark scores keep climbing, yet the same agent failures resu

IMR: Iterative Mode-World Weighted Regression for Multi-Agent Trajectory Prediction

Model ReleasesDGX agent

arXiv:2607.05705v1 Announce Type: cross Abstract: Multi-agent motion prediction is essential for automated vehicles to understand the intentions of surrounding vehicles. However, previous prediction-b

7 Jul 2026

A Reliable Context-Aware and Temporal Planning Framework for Autonomous Driving

Model ReleasesDGX agent

arXiv:2607.04689v1 Announce Type: cross Abstract: Safe operation of autonomous vehicles in dense urban traffic depends on perception and planning that remain reliable when onboard sensing is degraded.

A Technical Survey of Reinforcement Learning Techniques for Large Language Models

Model ReleasesDGX agent

arXiv:2507.04136v2 Announce Type: replace Abstract: This survey offers a comprehensive foundation on the integration of RL with language models, highlighting prominent algorithms such as Proximal Poli

Auditing the Audit: Five Failure Modes in Benchmark-Validity Audits

Model ReleasesDGX agent

arXiv:2607.02586v1 Announce Type: new Abstract: Governance frameworks ask AI providers and auditors for documented evaluation evidence, and perturbation-based construct-validity audits are a common fo

CTForensics: A Comprehensive Dataset and Method for AI-Generated CT Image Detection

Model ReleasesDGX agent

arXiv:2603.01878v2 Announce Type: replace Abstract: Recent advances in generative AI have made synthetic Computed Tomography (CT) images increasingly realistic, enabling promising applications in medi

Elastic Gang: Per-Token Membership Change for a Hard-Barriered LLM Inference Gang Co-Scheduled with OS Processes

Local AiDGX agent

arXiv:2607.04668v1 Announce Type: cross Abstract: On-device LLM decoding is a hard-barriered CPU-SIMD computation that wants every core for milliseconds per token, while the rest of the OS wants those

Embodied Operators and Benchmarking: Toward Reusable and Deployable Embodied Intelligence Systems

Model ReleasesDGX agent

arXiv:2607.03283v1 Announce Type: new Abstract: Embodied intelligence systems require not only end-to-end policy models, but also reusable functional modules that transform multimodal observations, ro

Evaluating Agentic Harness Systems for Autonomous Computational Pathology

Model ReleasesDGX agent

arXiv:2607.02598v1 Announce Type: new Abstract: Autonomous computational pathology (ACP) converts high-level pathology analysis goals into executable, traceable and clinically bounded workflows. Reali

Governed Individuation: Cryptographically Decoupling an Agent's Learning from Its Authority

Model ReleasesDGX agent

arXiv:2607.04613v1 Announce Type: new Abstract: Autonomous agents are moving from sandboxed text generators to operators of code, data, and physical infrastructure, and they increasingly learn while d

GPU-Accelerated Polygonal Signed Distance Functions for Real-Time Collision Avoidance

Local AiDGX agent

arXiv:2607.04310v1 Announce Type: new Abstract: Optimization-based local planning and control require high-rate collision-avoidance constraint evaluation over a prediction horizon. In obstacle-dense e

HAS-Bench: Evaluating LLM-Based Human-Agent Systems under Configurable Human Participation

Model ReleasesDGX agent

arXiv:2607.04329v1 Announce Type: new Abstract: Large language models increasingly operate in settings where humans are active collaborators rather than passive task providers. We introduce HAS-Framew

IndustryNav: Exploring Spatial Reasoning of Embodied Agents in Dynamic Industrial Navigation

Model ReleasesDGX agent

arXiv:2511.17384v2 Announce Type: replace-cross Abstract: While Visual Large Language Models (VLLMs) show great promise as embodied agents, they continue to face substantial challenges in spatial reas

iVISION-2DCD: A Long-Term Change Detection Dataset for Large-Scale Outdoor Construction Monitoring

Model ReleasesDGX agent

arXiv:2607.03553v1 Announce Type: new Abstract: Automation in construction is essential for reducing costs and human errors in large-scale projects. We approach the construction progress monitoring fr

Learning to Suppress SPAD-based LiDAR Flare

Model ReleasesDGX agent

arXiv:2607.03247v1 Announce Type: new Abstract: Single-Photon Avalanche Diode (SPAD)-based Light Detection and Ranging (LiDAR) is emerging for autonomous vehicles due to its high sensitivity and preci

Medi-Gemma: A Hybrid Clinical Decision Support System Integrating Deterministic EMR Analytics and Retrieval-Augmented Generation

Model ReleasesDGX agent

arXiv:2607.04907v1 Announce Type: new Abstract: Deploying Large Language Models (LLMs) in high-stakes clinical settings remains limited by structural hallucinations, weak deterministic reasoning over

Multi-Agent Reinforcement Learning for V2X Resource Allocation: Disentangling MARL Challenges Through Benchmarking

Model ReleasesDGX agent

arXiv:2603.06607v2 Announce Type: replace-cross Abstract: Radio resource allocation (RRA) is a critical function in cellular vehicle-to-everything (C-V2X) networks, where vehicles must share limited w

OpenGlass: A Sensing-Computing Split Architecture for Local MLLM-Driven Real-Time Visual Assistance

Local AiDGX agent

arXiv:2607.03213v1 Announce Type: cross Abstract: We present OpenGlass, an open-source, privacy-oriented, local-first system for low-latency multimodal visual assistance, with a primary focus on blind

RES-DARE: Failure-Aware Expert Adaptation and Rollback-Safe Self-Repair for Intrusion Detection

Model ReleasesDGX agent

arXiv:2607.02687v1 Announce Type: cross Abstract: Intrusion detection systems are often trained under static benchmark conditions, although deployed network environments are affected by traffic drift,

SovereignPA-Bench: Evaluating User-Owned Personal Agents under Evolving Intent, Platform Mediation, and Consent Constraints

Model ReleasesDGX agent

arXiv:2607.05363v1 Announce Type: new Abstract: Personal agents are becoming persistent user-owned intermediaries: they remember preferences, filter platform-mediated information, use tools, and negot

VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Model ReleasesDGX agent

arXiv:2407.11691v5 Announce Type: replace Abstract: We present VLMEvalKit: an open-source toolkit for evaluating large multi-modality models based on PyTorch. The toolkit aims to provide a user-friend

6 Jul 2026

Shift into high gear with agents: Securing the software-defined vehicle

Model ReleasesDGX agent

The automotive industry is at a pivotal crossroads as it hits the gas on adopting new technology. The era of the traditional connected vehicle has shifted into the age of the software-defined vehicle

3 Jul 2026

Assessing VLM Reliability for Medical Image Quality Evaluation Under Corruption and Bias

Model ReleasesDGX agent

arXiv:2607.01973v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are increasingly applied in medical tasks such as pathology description, report generation, and visual question answerin

Controllable Sim Agents with Behavior Latents

Model ReleasesDGX agent

arXiv:2607.02496v1 Announce Type: cross Abstract: Realistic traffic simulation requires agents that imitate logged behavior and can also be steered along interpretable axes. Such controllability enabl

IonSense-QKG: A Quantum-Readiness Metadata Framework for Lithium-Ion Battery Dataset Discovery

Model ReleasesDGX agent

arXiv:2607.01286v1 Announce Type: new Abstract: Public lithium-ion battery datasets are increasingly used for state-of-health estimation, remaining-useful-life prediction, anomaly detection, electroch

MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering

Model ReleasesDGX agent

arXiv:2607.01420v1 Announce Type: cross Abstract: As grounded QA systems are increasingly deployed in AI assistants, accurately attributing generated answers to evidence is critical for user trust and

Pre-Flight: A Benchmark for Evaluating Large Language Models on Aviation Operational Knowledge

Model ReleasesDGX agent

arXiv:2607.01829v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly proposed for aviation business operations, from documentation and training generation to customer facing a

Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness

Model ReleasesDGX agent

arXiv:2510.04484v2 Announce Type: replace-cross Abstract: The ability to control LLMs' emulated emotional states and personality traits is an essential step in enabling rich, human-centered interactio

Robust for the Wrong Reasons: The Representational Geometry of LLM Robustness to Science Skepticism

Model ReleasesDGX agent

arXiv:2607.01951v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly consulted on contested scientific questions, raising the concern that they will sycophantically retreat

← Previous
1…227228229230231…238
Next →