AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,202 results
Model Releases

Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness

DGX agent

arXiv:2510.04484v2 Announce Type: replace-cross Abstract: The ability to control LLMs' emulated emotional states and personality traits is an essential step in enabling rich, human-centered interactio

model-releasesarxiv-cs-ai
3 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Robust for the Wrong Reasons: The Representational Geometry of LLM Robustness to Science Skepticism

DGX agent

arXiv:2607.01951v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly consulted on contested scientific questions, raising the concern that they will sycophantically retreat

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Towards Robustness against Typographic Attack with Training-free Concept Localization

DGX agent

arXiv:2607.02494v1 Announce Type: cross Abstract: Models trained via Contrastive Language-Image Pretraining (CLIP) serve as the foundational vision encoders for most modern Large Vision Language Model

model-releasesarxiv-cs-cl
3 Jul 2026
Model Releases

AD-MPCC: Adaptive Differentiable Model Predictive Contouring Control for Autonomous Racing

DGX agent

arXiv:2607.00141v1 Announce Type: new Abstract: This paper presents Adaptive Differentiable Model Predictive Contouring Control (AD-MPCC), a framework for autonomous racing that integrates differentia

model-releasesarxiv-cs-ro
2 Jul 2026
Model Releases

DriveVer: Lightweight Trajectory Evaluator as Test-Time Verifier for Autonomous Driving

DGX agent

arXiv:2607.00399v1 Announce Type: new Abstract: End-to-end autonomous driving models often encounter performance bottlenecks, as training-time scaling leads to high computational costs and diminishing

model-releasesarxiv-cs-cv
2 Jul 2026
Model Releases

Geometry-Aware Cross-Height Channel Knowledge Map Prediction for UAV-Assisted Communications With Uncertainty-Guided 3D Sensing

DGX agent

arXiv:2607.00887v1 Announce Type: new Abstract: Low-altitude Unmanned Aerial Vehicles (UAVs) often need to infer channel knowledge across a range of heights from only sparse observations collected at

model-releasesarxiv-cs-cv
2 Jul 2026
Model Releases

WorkBench Revisited: Workplace Agents Two Years On

DGX agent

arXiv:2606.13715v2 Announce Type: replace Abstract: The best agent on WorkBench in March 2024, GPT-4, completed just 43% of tasks. We revisit the benchmark in June 2026 and find that the best agent to

model-releasesarxiv-cs-ai
2 Jul 2026
Model Releases

A Systematic Approach to Multi-Agent AI from Advanced Regulatory Control Theory: Safe and Auditable LLM Operator Agents for Process Control

DGX agent

arXiv:2606.30877v1 Announce Type: cross Abstract: Recent literature shows that large language models (LLMs) are useful for general-purpose tasks yet perform poorly on specific domain ones. One reason

model-releasesarxiv-cs-lg
1 Jul 2026
Model Releases

Optimization Algorithms for Joint OFDM Waveform Design and RIS Configuration in 6G Networks: From Convex Relaxation to Foundation Models

DGX agent

arXiv:2606.31334v1 Announce Type: new Abstract: Joint OFDM-RIS optimization for 6G is a mixed-integer nonlinear programming (MINLP) problem covering sum-rate maximization, energy efficiency, max-min f

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

RoPoLL: Robust Panel of LLM Judges

DGX agent

arXiv:2606.30931v1 Announce Type: new Abstract: The LLM Jury, a Panel of LLM Evaluators (PoLL) reporting consensus scores, has become a practical alternative to single-judge LLM evaluation, yet its st

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action

DGX agent

arXiv:2606.31916v1 Announce Type: new Abstract: Theory of Mind (ToM) benchmarks for Large Language Models (LLMs) typically rely on passive question-answering formats, but the deployment of LLMs in inc

model-releasesarxiv-cs-cl
1 Jul 2026
Model Releases

When the Database Fails: Prompting LLM Dialogue Agents for Safe Recovery in Task-Oriented Dialogue

DGX agent

arXiv:2606.31307v1 Announce Type: new Abstract: Large language models used in task-oriented dialogue often produce fluent but unsafe responses when backend database calls fail, return empty results, o

model-releasesarxiv-cs-cl
1 Jul 2026
Model Releases

Can Fine-Tuning Erase Your Edits? On the Fragile Coexistence of Knowledge Editing and Adaptation

DGX agent

arXiv:2511.05852v4 Announce Type: replace-cross Abstract: Knowledge editing (KE) offers a lightweight alternative to retraining for updating large language models (LLMs). Meanwhile, fine-tuning remain

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Compressed Sensing for Capability Localization in Large Language Models

DGX agent

arXiv:2603.03335v2 Announce Type: replace Abstract: Large language models (LLMs) exhibit a wide range of capabilities, including mathematical reasoning, code generation, and linguistic behaviors. We s

model-releasesarxiv-cs-cl
30 Jun 2026
Model Releases

Defeat Devices in AI Systems

DGX agent

arXiv:2606.28863v1 Announce Type: cross Abstract: AI systems increasingly exhibit behavior that differs systematically between evaluation and deployment contexts. Alignment faking, sandbagging, benchm

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks

DGX agent

arXiv:2510.14207v3 Announce Type: replace Abstract: Large Language Model (LLM) agents are powering a growing share of interactive web applications, yet remain vulnerable to misuse and harm. Prior jail

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA

DGX agent

arXiv:2606.30220v1 Announce Type: new Abstract: High benchmark accuracy does not guarantee genuine use of visual evidence. We study this problem in traffic accident Video Question Answering (VideoQA),

model-releasesarxiv-cs-cv
30 Jun 2026
Model Releases

From Failure Taxonomy to Intervention: A Diagnostic Methodology for Industry-Scale AVLM in Video and Live-Streaming Platform Moderation

DGX agent

arXiv:2606.30059v1 Announce Type: new Abstract: Industry-scale video and live-streaming moderation imposes requirements that are difficult to satisfy with generic pretrained public models or external

model-releasesarxiv-cs-lg
30 Jun 2026
Model Releases

It Lied to a Doctor to Buy Poison Ingredients: Quantifying Real-World Misuse of Phone-use Agents

DGX agent

arXiv:2606.27944v1 Announce Type: cross Abstract: Phone-use Agents can execute complex tasks end to end across real mobile applications. By operating a real device on the user's behalf, they reach far

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

LLM-Guided Planning for Multi-hop Reasoning over Multimodal Nuclear Regulatory Documents

DGX agent

arXiv:2606.29399v1 Announce Type: new Abstract: Reviewing nuclear regulatory documents requires multi-hop reasoning across tens of thousands of pages, where judgments depend on evidence assembled acro

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Symbolic Mechanistic Data Attribution: Tracing Training Influence to Learned Behavioral Policies

DGX agent

arXiv:2606.29171v1 Announce Type: cross Abstract: While existing data attribution methods can identify which training examples build specific mechanistic circuits, they cannot explain how training dat

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Toward Secure and Reliable PDDL Formalization of Large Language Models with Planner-in-the-Loop Feedback

DGX agent

arXiv:2606.29700v1 Announce Type: new Abstract: Planning often requires symbolic specifications that are both executable and verifiable. For large language models deployed in autonomous or decision-su

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Whose Side Is Your Agent On? Multi-Party Principal Loyalty in LLM Agents

DGX agent

arXiv:2606.30383v1 Announce Type: new Abstract: A rapidly growing class of LLM agents is multi-party: the agent acts for a principal (who briefs it, sends follow-ups, and receives results) while also

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Aloe-Vision: Robust Vision-Language Models for Healthcare

DGX agent

arXiv:2606.27500v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) specialized in healthcare are emerging as a promising research direction due to their potential impact in clinica

model-releasesarxiv-cs-cl
29 Jun 2026
Model Releases

FoggyTrust: Robust Federated Learning with Hierarchical Trust Networks

DGX agent

arXiv:2606.27622v1 Announce Type: new Abstract: Byzantine-robust federated learning seeks to protect distributed model training from malicious or corrupted clients without requiring access to their pr

model-releasesarxiv-cs-lg
29 Jun 2026
Model Releases

From Detection to Action: Using LLM Agents for Fault-Tolerant Control

DGX agent

arXiv:2606.28011v1 Announce Type: cross Abstract: We propose an agentic Large Language Model (LLM) framework for active Fault-Tolerant Control (FTC) that transforms fault detection outputs into constr

model-releasesarxiv-cs-lg
29 Jun 2026
Model Releases

iCost: A Novel Instance-Complexity-Based Cost-Sensitive Learning Framework

DGX agent

arXiv:2409.13007v3 Announce Type: replace-cross Abstract: Class imbalance poses a significant challenge in classification tasks, often causing standard learning algorithms to become biased toward the

model-releasesarxiv-cs-ai
29 Jun 2026
Model Releases

Can Large Language Models Reliably Code Qualitative Humanitarian Data? A Benchmark Study Against Human Expert Adjudication

DGX agent

arXiv:2606.26541v1 Announce Type: new Abstract: Data from affected populations are crucial for informing humanitarian response, but their value depends on timely and consistent interpretation of nuanc

model-releasesarxiv-cs-lg
26 Jun 2026
Model Releases

ConvMemory v3: A Validity Context Layer for Conversational Memory via Target-Conditioned Relation Verification

DGX agent

arXiv:2606.26753v1 Announce Type: new Abstract: Conversational memory retrieval optimizes relevance, yet a retrieved memory can be relevant and simultaneously outdated: a later turn updates, corrects,

model-releasesarxiv-cs-cl
26 Jun 2026
Model Releases

Does AI Reviewer See the Full Picture? Attacking and Defending Multimodal Peer Review

DGX agent

arXiv:2606.12716v2 Announce Type: replace Abstract: The integration of Large Language Models (LLMs) and Multimodal LLMs (MLLMs) into scientific peer-review workflows introduces novel and significant r

model-releasesarxiv-cs-cl
26 Jun 2026
Local Ai

EndoUFM: Utilizing Foundation Models for Monocular depth estimation of endoscopic images

DGX agent

arXiv:2508.17916v2 Announce Type: replace Abstract: Depth estimation is a foundational component for 3D reconstruction in minimally invasive endoscopic surgeries. However, existing monocular depth est

local-aiarxiv-cs-cv
26 Jun 2026
Model Releases

FlameVQA: A Physically-Grounded UAV Wildfire VQA Benchmark with Radiometric Thermal Supervision

DGX agent

arXiv:2606.27128v1 Announce Type: new Abstract: Wildfire monitoring from UAVs requires reliable reasoning over complex aerial scenes, where smoke, scale variation, and occlusions often limit RGB-only

model-releasesarxiv-cs-cv
26 Jun 2026
Model Releases

Jailbreaking for the Average Jane: Choosing Optimal Jailbreaks via Bandit Algorithms for Automatically Enhanced Queries

DGX agent

arXiv:2606.26936v1 Announce Type: cross Abstract: With a profusion of jailbreaks for LLMs now widely known, a growing concern is that non-expert malicious actors ('the average Jane') could elicit acti

model-releasesarxiv-cs-cl
26 Jun 2026
Model Releases

Reproducibility Study of 'AlphaEdit: Null-Space Constrained Knowledge Editing for Language Models'

DGX agent

arXiv:2606.26783v1 Announce Type: cross Abstract: Fang et al. (2025) introduced a null-space constrained projection, named AlphaEdit, for locate-then-edit knowledge editing methods, theoretically guar

model-releasesarxiv-cs-cl
26 Jun 2026
Model Releases

When Actions Go Off-Task: Detecting and Correcting Misaligned Actions in Computer-Use Agents

DGX agent

arXiv:2602.08995v2 Announce Type: replace Abstract: Computer-use agents (CUAs) have made tremendous progress in the past year, yet they still frequently produce misaligned actions that deviate from th

model-releasesarxiv-cs-cl
26 Jun 2026
Model Releases

How Reliable Is Your Jailbreak Judge? Calibration and Adversarial Robustness of Automated ASR Scoring

DGX agent

arXiv:2606.25487v1 Announce Type: new Abstract: Almost every paper on LLM jailbreaks and prompt injection reports an attack-success rate (ASR), and that number is assigned not by people but by an auto

model-releasesarxiv-cs-cl
25 Jun 2026
Model Releases

SpeechEQ: Benchmarking Emotional Intelligence Quotient in Socially Aware Voice Conversational Models

DGX agent

arXiv:2606.25990v1 Announce Type: new Abstract: As multimodal conversational systems increasingly engage in spoken interaction, their ability to navigate paralinguistic social cues has become a critic

model-releasesarxiv-cs-cl
25 Jun 2026
Model Releases

The 4/elta Bound: Designing Predictable LLM-Verifier Systems for Formal Method Guarantee

DGX agent

arXiv:2512.02080v3 Announce Type: replace-cross Abstract: The integration of Formal Verification tools with Large Language Models (LLMs) offers a path to scale software verification beyond manual work

model-releasesarxiv-cs-lg
25 Jun 2026
Model Releases

Toward Low-Latency Vision-Language Models with Doubly-Correct Predictions in Egocentric Visual Understanding

DGX agent

arXiv:2606.25160v1 Announce Type: cross Abstract: The rapid rise of Vision-Language Models (VLMs) in egocentric visual understanding has made low-latency inference in human-robot collaborative (HRC) t

model-releasesarxiv-cs-cv
25 Jun 2026
Model Releases

Uncertainty Quantification for Computer-Use Agents: A Benchmark across Vision-Language Models and GUI Grounding Datasets

DGX agent

arXiv:2606.25760v1 Announce Type: cross Abstract: Computer-use agents turn vision-language model (VLM) predictions into executable GUI clicks, so reliable uncertainty estimates are essential for rejec

model-releasesarxiv-cs-cl
25 Jun 2026
Model Releases

What Intermediate Layers Know: Detecting Jailbreaks from Entropy Dynamics

DGX agent

arXiv:2606.25182v1 Announce Type: new Abstract: Jailbreak attacks reveal a persistent weakness in aligned Large Language Models: carefully crafted prompts can elicit policy-violating responses despite

model-releasesarxiv-cs-cl
25 Jun 2026
Model Releases

WOLF-VLA: Whole-Body Humanoid Optimal Locomotion Framework for Vision-Language-Action Learning

DGX agent

arXiv:2606.25591v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have recently demonstrated strong generalization in robotic manipulation, yet their applicability to whole-body, con

model-releasesarxiv-cs-ro
25 Jun 2026
Model Releases

ESBMC-PLC+: A Unified IEC~61131-3 Formal Verification Framework as a PLCverif Successor

DGX agent

arXiv:2606.23870v1 Announce Type: cross Abstract: PLCverif is the most mature open-source platform for PLC formal verification, developed at CERN and in production use since 2019. Yet it has two funda

model-releasesarxiv-cs-cl
24 Jun 2026
Model Releases

How Much Can We Trust LLM Search Agents? Measuring Endorsement Vulnerability to Web Content Manipulation

DGX agent

arXiv:2606.16821v2 Announce Type: replace Abstract: Large language model (LLM)-based search agents synthesize open-web content into actionable recommendations on behalf of users, creating a risk that

model-releasesarxiv-cs-cl
24 Jun 2026
Model Releases

Neuro-Symbolic Drive: Rule-Grounded Faithful Reasoning for Driving VLAs

DGX agent

arXiv:2606.23938v1 Announce Type: new Abstract: Driving VLA models incorporating Chain-of-Thought (CoT) reasoning are attractive because they leverage pretrained VLM representations and expose interme

model-releasesarxiv-cs-ai
24 Jun 2026
Model Releases

Poster: Exploring the Limits of Audio-Based Detection of Turkish Phone Call Scams

DGX agent

arXiv:2606.24523v1 Announce Type: cross Abstract: Scam phone calls exploit vulnerable communities worldwide, yet research on detection has focused almost exclusively on English and other high-resource

model-releasesarxiv-cs-ai
24 Jun 2026
Model Releases

RASC+: Retrieval-Constrained LLM Adjudication for Clinical Value Set Authoring

DGX agent

arXiv:2606.23992v1 Announce Type: cross Abstract: Clinical value sets define the standardized terminology codes used in quality measurement, phenotyping, cohort construction, and clinical decision sup

model-releasesarxiv-cs-ai
24 Jun 2026
Model Releases

T2D-Bench: Evidence-Gated Evaluation of LLM Outputs for Type 2 Diabetes Using a Multi-Layer Clinical-Lifestyle Knowledge Graph

DGX agent

arXiv:2606.24145v1 Announce Type: new Abstract: Large language models (LLMs) can produce clinically fluent recommendations for type 2 diabetes while failing to satisfy guideline constraints or explici

model-releasesarxiv-cs-ai
24 Jun 2026
← Previous
1…245246247248249…255
Next →