AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,202 results
Model Releases

HRIBench: Benchmarking Interaction-Centric Human-Robot Collaboration

DGX agent

arXiv:2607.13056v1 Announce Type: cross Abstract: Current vision-language-action (VLA) benchmarks primarily evaluate isolated manipulation skills while leaving human-robot interaction structure largel

model-releasesarxiv-cs-lg
16 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation

DGX agent

arXiv:2607.09142v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in online medical consultation, yet existing benchmarks remain poorly aligned with real clini

model-releasesarxiv-cs-ai
16 Jul 2026
Model Releases

Not All Needles Are Found: How Fact Distribution and Don't Make It Up Prompts Shape Retrieval, Reasoning, and Hallucination in Long-Context LLMs

DGX agent

arXiv:2601.02023v2 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) increasingly utilize massive context windows as working memory for autonomous tasks, their reliability fluctua

model-releasesarxiv-cs-ai
16 Jul 2026
Model Releases

nuTruck: Benchmarking Autonomous Driving Planning for Distributed Electric-drive Trucks

DGX agent

arXiv:2607.13704v1 Announce Type: new Abstract: The dominance of traditional rule-based methods in autonomous driving has gradually been replaced by learning-based approaches. While learning-based pla

model-releasesarxiv-cs-ro
16 Jul 2026
Model Releases

Securing LLMs in the Wild: Privacy and Security Challenges at the Edge

DGX agent

arXiv:2607.13088v1 Announce Type: cross Abstract: Large Language Models (LLMs) are rapidly moving from research settings into the wild, deployed on enterprise infrastructure, personal devices, and edg

model-releasesarxiv-cs-lg
16 Jul 2026
Model Releases

Agents Don't Just Agree, They Remember: Benchmarking Persistent Sycophancy in Stateful Personal Agents

DGX agent

arXiv:2607.10526v2 Announce Type: replace Abstract: Stateful personal agents increasingly maintain long-term user profiles, episodic memories, and reusable skills. This persistence turns conversationa

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Are we Merging the Right Models? Impact of Expert Training Duration on Model Merging for LLMs

DGX agent

arXiv:2607.11997v1 Announce Type: cross Abstract: Multi-task model merging combines separately trained expert models into a single model that handles all tasks without co-training. Standard practice m

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Breaking Deja Vu: Independent Auditing of Visual Place Recognition through Vision-Language Reasoning

DGX agent

arXiv:2607.12818v1 Announce Type: new Abstract: Visual place recognition (VPR) is a key enabler of accurate localization and long-term autonomous navigation in robotics applications, such as loop clos

model-releasesarxiv-cs-cv
15 Jul 2026
Model Releases

Continual Learning with Elastic Regularization and Synthetic Replay for Federated MLLM Fine-Tuning

DGX agent

arXiv:2607.12112v1 Announce Type: cross Abstract: Federated fine-tuning of Multimodal Large Language Models (MLLMs) across distributed networks enables privacy-sensitive adaptation to evolving data st

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Epistemic Stance Flexibility Probing: Measuring Prompt-Conditioned Register Shift in Large Language Models

DGX agent

arXiv:2607.12739v1 Announce Type: new Abstract: A language model may be asked either what experts believe about a contested claim or what it believes about the claim itself. A trustworthy conversation

model-releasesarxiv-cs-cl
15 Jul 2026
Model Releases

How Inference Compute Shapes Frontier LLM Evaluation

DGX agent

arXiv:2606.17930v2 Announce Type: replace Abstract: AI evaluations are shifting toward harder tasks that benefit from longer trajectories involving tool use and iterative problem solving. As a result,

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

How Many Tasks Are Enough for Agent Benchmark Decisions? A Replay Analysis of Public LLM Agent Benchmarks

DGX agent

arXiv:2607.12338v1 Announce Type: new Abstract: Agent benchmarks often compare two agents after all tasks have run, but costly evaluations make partial runs tempting. A task fraction alone does not sh

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

How to Realize Recursively Self-Improving Agents and Personal Singularity: A Goal-, Scope-, Tool-, and Benchmark-Driven Multi-Agent Architecture

DGX agent

arXiv:2607.12254v1 Announce Type: new Abstract: Large language model (LLM) agents can increasingly plan, use tools, maintain memory, and execute long-horizon tasks. These advances motivate two linked

model-releasesarxiv-cs-cv
15 Jul 2026
Model Releases

Knowledge- and Gradient-Guided Reinforcement Learning for Parametrized Action Markov Decision Processes

DGX agent

arXiv:2607.12924v1 Announce Type: new Abstract: In this paper, we study Reinforcement Learning in Parametrized Action Markov Decision Processes (PAMDP), where each decision consists of a symbolic acti

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Ontology-Amplified Distillation and Contextuality Auditing for Sovereign Enterprise Language Models: A Combined Proof-of-Mechanism and Negative-Results Method Study

DGX agent

arXiv:2607.11948v1 Announce Type: new Abstract: Regulated financial institutions operating under data-residency rules need tenant-owned language models that can run inside the institution's perimeter.

model-releasesarxiv-cs-ai
15 Jul 2026
Local Ai

OmniOPSD: Rationale-Privileged On-Policy Self-Distillation for Affective Computing

DGX agent

arXiv:2606.15920v2 Announce Type: replace Abstract: Reinforcement learning for multimodal large language models (MLLMs) is often hindered by severe reward sparsity in complex reasoning tasks. This cha

local-aiarxiv-cs-cv
10 Jul 2026
Model Releases

Secure Decentralized Federated Learning via Gossip and Virtual Voting

DGX agent

arXiv:2607.08651v1 Announce Type: new Abstract: Decentralized federated learning (DFL) removes the central server by letting nodes exchange model updates through peer-to-peer gossip, but existing goss

model-releasesarxiv-cs-lg
10 Jul 2026
Model Releases

Shift & Drift: A Zero-Shot Benchmark for Generalizable and Robust Autonomous Driving Motion Planning

DGX agent

arXiv:2607.07844v1 Announce Type: cross Abstract: While closed-loop motion planners trained on large-scale, object-level datasets, e.g., nuPlan, demonstrate strong in-distribution (ID) performance, th

model-releasesarxiv-cs-ai
10 Jul 2026
Model Releases

Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI Agents

DGX agent

arXiv:2607.07474v1 Announce Type: cross Abstract: Agentic red-teaming benchmarks report whether an injected agent was compromised as a single bit: the attack succeeded, or it did not. We argue that th

model-releasesarxiv-cs-ai
9 Jul 2026
Model Releases

Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages

DGX agent

arXiv:2607.06596v1 Announce Type: cross Abstract: Trusted monitoring is a central defense in AI control: a cheaper trusted model scores an untrusted model's actions for sabotage, and the most suspicio

model-releasesarxiv-cs-lg
9 Jul 2026
Model Releases

Multi-Agent Robotic Control with Onboard Vision-Language Models

DGX agent

arXiv:2607.07403v1 Announce Type: cross Abstract: Vision Language Models (VLMs) and Vision Language Action (VLA) models have shown promise in robotic control. Yet, they face significant challenges reg

model-releasesarxiv-cs-ro
9 Jul 2026
Model Releases

SmartHomeSecure: Automated Detection and Repair of Smart Home Configuration Errors Using Large Language Models

DGX agent

arXiv:2607.06748v1 Announce Type: cross Abstract: Smart home automation platforms increasingly rely on user-authored YAML configuration files to define device behaviors, but these files are prone to s

model-releasesarxiv-cs-ai
9 Jul 2026
Model Releases

aiAuthZ: Off-Host, Identity-Bound Authorization for AI Agents

DGX agent

arXiv:2607.05518v1 Announce Type: cross Abstract: AI agents issue tool calls on the basis of text they cannot verify, so any party who controls part of the context can forge the appearance of authorit

model-releasesarxiv-cs-ai
8 Jul 2026
Model Releases

IMR: Iterative Mode-World Weighted Regression for Multi-Agent Trajectory Prediction

DGX agent

arXiv:2607.05705v1 Announce Type: cross Abstract: Multi-agent motion prediction is essential for automated vehicles to understand the intentions of surrounding vehicles. However, previous prediction-b

model-releasesarxiv-cs-ai
8 Jul 2026
Model Releases

A Reliable Context-Aware and Temporal Planning Framework for Autonomous Driving

DGX agent

arXiv:2607.04689v1 Announce Type: cross Abstract: Safe operation of autonomous vehicles in dense urban traffic depends on perception and planning that remain reliable when onboard sensing is degraded.

model-releasesarxiv-cs-cv
7 Jul 2026
Model Releases

A Technical Survey of Reinforcement Learning Techniques for Large Language Models

DGX agent

arXiv:2507.04136v2 Announce Type: replace Abstract: This survey offers a comprehensive foundation on the integration of RL with language models, highlighting prominent algorithms such as Proximal Poli

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

Auditing the Audit: Five Failure Modes in Benchmark-Validity Audits

DGX agent

arXiv:2607.02586v1 Announce Type: new Abstract: Governance frameworks ask AI providers and auditors for documented evaluation evidence, and perturbation-based construct-validity audits are a common fo

model-releasesarxiv-cs-lg
7 Jul 2026
Model Releases

CTForensics: A Comprehensive Dataset and Method for AI-Generated CT Image Detection

DGX agent

arXiv:2603.01878v2 Announce Type: replace Abstract: Recent advances in generative AI have made synthetic Computed Tomography (CT) images increasingly realistic, enabling promising applications in medi

model-releasesarxiv-cs-cv
7 Jul 2026
Local Ai

Elastic Gang: Per-Token Membership Change for a Hard-Barriered LLM Inference Gang Co-Scheduled with OS Processes

DGX agent

arXiv:2607.04668v1 Announce Type: cross Abstract: On-device LLM decoding is a hard-barriered CPU-SIMD computation that wants every core for milliseconds per token, while the rest of the OS wants those

local-aiarxiv-cs-ai
7 Jul 2026
Model Releases

Embodied Operators and Benchmarking: Toward Reusable and Deployable Embodied Intelligence Systems

DGX agent

arXiv:2607.03283v1 Announce Type: new Abstract: Embodied intelligence systems require not only end-to-end policy models, but also reusable functional modules that transform multimodal observations, ro

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

Evaluating Agentic Harness Systems for Autonomous Computational Pathology

DGX agent

arXiv:2607.02598v1 Announce Type: new Abstract: Autonomous computational pathology (ACP) converts high-level pathology analysis goals into executable, traceable and clinically bounded workflows. Reali

model-releasesarxiv-cs-cv
7 Jul 2026
Model Releases

Governed Individuation: Cryptographically Decoupling an Agent's Learning from Its Authority

DGX agent

arXiv:2607.04613v1 Announce Type: new Abstract: Autonomous agents are moving from sandboxed text generators to operators of code, data, and physical infrastructure, and they increasingly learn while d

model-releasesarxiv-cs-ai
7 Jul 2026
Local Ai

GPU-Accelerated Polygonal Signed Distance Functions for Real-Time Collision Avoidance

DGX agent

arXiv:2607.04310v1 Announce Type: new Abstract: Optimization-based local planning and control require high-rate collision-avoidance constraint evaluation over a prediction horizon. In obstacle-dense e

local-aiarxiv-cs-ro
7 Jul 2026
Model Releases

HAS-Bench: Evaluating LLM-Based Human-Agent Systems under Configurable Human Participation

DGX agent

arXiv:2607.04329v1 Announce Type: new Abstract: Large language models increasingly operate in settings where humans are active collaborators rather than passive task providers. We introduce HAS-Framew

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

IndustryNav: Exploring Spatial Reasoning of Embodied Agents in Dynamic Industrial Navigation

DGX agent

arXiv:2511.17384v2 Announce Type: replace-cross Abstract: While Visual Large Language Models (VLLMs) show great promise as embodied agents, they continue to face substantial challenges in spatial reas

model-releasesarxiv-cs-cv
7 Jul 2026
Model Releases

iVISION-2DCD: A Long-Term Change Detection Dataset for Large-Scale Outdoor Construction Monitoring

DGX agent

arXiv:2607.03553v1 Announce Type: new Abstract: Automation in construction is essential for reducing costs and human errors in large-scale projects. We approach the construction progress monitoring fr

model-releasesarxiv-cs-cv
7 Jul 2026
Model Releases

Learning to Suppress SPAD-based LiDAR Flare

DGX agent

arXiv:2607.03247v1 Announce Type: new Abstract: Single-Photon Avalanche Diode (SPAD)-based Light Detection and Ranging (LiDAR) is emerging for autonomous vehicles due to its high sensitivity and preci

model-releasesarxiv-cs-cv
7 Jul 2026
Model Releases

Medi-Gemma: A Hybrid Clinical Decision Support System Integrating Deterministic EMR Analytics and Retrieval-Augmented Generation

DGX agent

arXiv:2607.04907v1 Announce Type: new Abstract: Deploying Large Language Models (LLMs) in high-stakes clinical settings remains limited by structural hallucinations, weak deterministic reasoning over

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

Multi-Agent Reinforcement Learning for V2X Resource Allocation: Disentangling MARL Challenges Through Benchmarking

DGX agent

arXiv:2603.06607v2 Announce Type: replace-cross Abstract: Radio resource allocation (RRA) is a critical function in cellular vehicle-to-everything (C-V2X) networks, where vehicles must share limited w

model-releasesarxiv-cs-ai
7 Jul 2026
Local Ai

OpenGlass: A Sensing-Computing Split Architecture for Local MLLM-Driven Real-Time Visual Assistance

DGX agent

arXiv:2607.03213v1 Announce Type: cross Abstract: We present OpenGlass, an open-source, privacy-oriented, local-first system for low-latency multimodal visual assistance, with a primary focus on blind

local-aiarxiv-cs-ai
7 Jul 2026
Model Releases

RES-DARE: Failure-Aware Expert Adaptation and Rollback-Safe Self-Repair for Intrusion Detection

DGX agent

arXiv:2607.02687v1 Announce Type: cross Abstract: Intrusion detection systems are often trained under static benchmark conditions, although deployed network environments are affected by traffic drift,

model-releasesarxiv-cs-lg
7 Jul 2026
Model Releases

SovereignPA-Bench: Evaluating User-Owned Personal Agents under Evolving Intent, Platform Mediation, and Consent Constraints

DGX agent

arXiv:2607.05363v1 Announce Type: new Abstract: Personal agents are becoming persistent user-owned intermediaries: they remember preferences, filter platform-mediated information, use tools, and negot

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

DGX agent

arXiv:2407.11691v5 Announce Type: replace Abstract: We present VLMEvalKit: an open-source toolkit for evaluating large multi-modality models based on PyTorch. The toolkit aims to provide a user-friend

model-releasesarxiv-cs-cv
7 Jul 2026
Model Releases

Assessing VLM Reliability for Medical Image Quality Evaluation Under Corruption and Bias

DGX agent

arXiv:2607.01973v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are increasingly applied in medical tasks such as pathology description, report generation, and visual question answerin

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Controllable Sim Agents with Behavior Latents

DGX agent

arXiv:2607.02496v1 Announce Type: cross Abstract: Realistic traffic simulation requires agents that imitate logged behavior and can also be steered along interpretable axes. Such controllability enabl

model-releasesarxiv-cs-lg
3 Jul 2026
Model Releases

IonSense-QKG: A Quantum-Readiness Metadata Framework for Lithium-Ion Battery Dataset Discovery

DGX agent

arXiv:2607.01286v1 Announce Type: new Abstract: Public lithium-ion battery datasets are increasingly used for state-of-health estimation, remaining-useful-life prediction, anomaly detection, electroch

model-releasesarxiv-cs-lg
3 Jul 2026
Model Releases

MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering

DGX agent

arXiv:2607.01420v1 Announce Type: cross Abstract: As grounded QA systems are increasingly deployed in AI assistants, accurately attributing generated answers to evidence is critical for user trust and

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Pre-Flight: A Benchmark for Evaluating Large Language Models on Aviation Operational Knowledge

DGX agent

arXiv:2607.01829v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly proposed for aviation business operations, from documentation and training generation to customer facing a

model-releasesarxiv-cs-ai
3 Jul 2026
← Previous
1…244245246247248…255
Next →