AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,237 results
Local Ai

OpenGlass: A Sensing-Computing Split Architecture for Local MLLM-Driven Real-Time Visual Assistance

DGX agent

arXiv:2607.03213v1 Announce Type: cross Abstract: We present OpenGlass, an open-source, privacy-oriented, local-first system for low-latency multimodal visual assistance, with a primary focus on blind

local-aiarxiv-cs-ai
7 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

RES-DARE: Failure-Aware Expert Adaptation and Rollback-Safe Self-Repair for Intrusion Detection

DGX agent

arXiv:2607.02687v1 Announce Type: cross Abstract: Intrusion detection systems are often trained under static benchmark conditions, although deployed network environments are affected by traffic drift,

model-releasesarxiv-cs-lg
7 Jul 2026
Model Releases

SovereignPA-Bench: Evaluating User-Owned Personal Agents under Evolving Intent, Platform Mediation, and Consent Constraints

DGX agent

arXiv:2607.05363v1 Announce Type: new Abstract: Personal agents are becoming persistent user-owned intermediaries: they remember preferences, filter platform-mediated information, use tools, and negot

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

DGX agent

arXiv:2407.11691v5 Announce Type: replace Abstract: We present VLMEvalKit: an open-source toolkit for evaluating large multi-modality models based on PyTorch. The toolkit aims to provide a user-friend

model-releasesarxiv-cs-cv
7 Jul 2026
Model Releases

Shift into high gear with agents: Securing the software-defined vehicle

DGX agent

The automotive industry is at a pivotal crossroads as it hits the gas on adopting new technology. The era of the traditional connected vehicle has shifted into the age of the software-defined vehicle

model-releasesgoogle-cloud-ai
6 Jul 2026
Model Releases

Assessing VLM Reliability for Medical Image Quality Evaluation Under Corruption and Bias

DGX agent

arXiv:2607.01973v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are increasingly applied in medical tasks such as pathology description, report generation, and visual question answerin

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Controllable Sim Agents with Behavior Latents

DGX agent

arXiv:2607.02496v1 Announce Type: cross Abstract: Realistic traffic simulation requires agents that imitate logged behavior and can also be steered along interpretable axes. Such controllability enabl

model-releasesarxiv-cs-lg
3 Jul 2026
Model Releases

IonSense-QKG: A Quantum-Readiness Metadata Framework for Lithium-Ion Battery Dataset Discovery

DGX agent

arXiv:2607.01286v1 Announce Type: new Abstract: Public lithium-ion battery datasets are increasingly used for state-of-health estimation, remaining-useful-life prediction, anomaly detection, electroch

model-releasesarxiv-cs-lg
3 Jul 2026
Model Releases

MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering

DGX agent

arXiv:2607.01420v1 Announce Type: cross Abstract: As grounded QA systems are increasingly deployed in AI assistants, accurately attributing generated answers to evidence is critical for user trust and

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Pre-Flight: A Benchmark for Evaluating Large Language Models on Aviation Operational Knowledge

DGX agent

arXiv:2607.01829v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly proposed for aviation business operations, from documentation and training generation to customer facing a

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness

DGX agent

arXiv:2510.04484v2 Announce Type: replace-cross Abstract: The ability to control LLMs' emulated emotional states and personality traits is an essential step in enabling rich, human-centered interactio

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Robust for the Wrong Reasons: The Representational Geometry of LLM Robustness to Science Skepticism

DGX agent

arXiv:2607.01951v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly consulted on contested scientific questions, raising the concern that they will sycophantically retreat

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Towards Robustness against Typographic Attack with Training-free Concept Localization

DGX agent

arXiv:2607.02494v1 Announce Type: cross Abstract: Models trained via Contrastive Language-Image Pretraining (CLIP) serve as the foundational vision encoders for most modern Large Vision Language Model

model-releasesarxiv-cs-cl
3 Jul 2026
Model Releases

AD-MPCC: Adaptive Differentiable Model Predictive Contouring Control for Autonomous Racing

DGX agent

arXiv:2607.00141v1 Announce Type: new Abstract: This paper presents Adaptive Differentiable Model Predictive Contouring Control (AD-MPCC), a framework for autonomous racing that integrates differentia

model-releasesarxiv-cs-ro
2 Jul 2026
Model Releases

Another fascinating paper on LLM Judges. (bookmark it) It's from Amazon, and they show that if you run panels of LLM judges, averaging their…

DGX agent

Another fascinating paper on LLM Judges. (bookmark it) It's from Amazon, and they show that if you run panels of LLM judges, averaging their scores is a trap. 'Overall, we establish that robust aggreg

model-releasesdair-ai--x
2 Jul 2026
Model Releases

DriveVer: Lightweight Trajectory Evaluator as Test-Time Verifier for Autonomous Driving

DGX agent

arXiv:2607.00399v1 Announce Type: new Abstract: End-to-end autonomous driving models often encounter performance bottlenecks, as training-time scaling leads to high computational costs and diminishing

model-releasesarxiv-cs-cv
2 Jul 2026
Model Releases

Geometry-Aware Cross-Height Channel Knowledge Map Prediction for UAV-Assisted Communications With Uncertainty-Guided 3D Sensing

DGX agent

arXiv:2607.00887v1 Announce Type: new Abstract: Low-altitude Unmanned Aerial Vehicles (UAVs) often need to infer channel knowledge across a range of heights from only sparse observations collected at

model-releasesarxiv-cs-cv
2 Jul 2026
Model Releases

WorkBench Revisited: Workplace Agents Two Years On

DGX agent

arXiv:2606.13715v2 Announce Type: replace Abstract: The best agent on WorkBench in March 2024, GPT-4, completed just 43% of tasks. We revisit the benchmark in June 2026 and find that the best agent to

model-releasesarxiv-cs-ai
2 Jul 2026
Model Releases

A Systematic Approach to Multi-Agent AI from Advanced Regulatory Control Theory: Safe and Auditable LLM Operator Agents for Process Control

DGX agent

arXiv:2606.30877v1 Announce Type: cross Abstract: Recent literature shows that large language models (LLMs) are useful for general-purpose tasks yet perform poorly on specific domain ones. One reason

model-releasesarxiv-cs-lg
1 Jul 2026
Model Releases

Loved the chat between @trq212 @_catwu @simonw at AI Eng summit. My top 13 takeaways from their session -> 1. Engineers should become better…

DGX agent

Loved the chat between @trq212 @_catwu @simonw at AI Eng summit. My top 13 takeaways from their session -> 1. Engineers should become better at product/business sense. 2. Don't worry about major rewri

model-releasesswyx--x
1 Jul 2026
Model Releases

Optimization Algorithms for Joint OFDM Waveform Design and RIS Configuration in 6G Networks: From Convex Relaxation to Foundation Models

DGX agent

arXiv:2606.31334v1 Announce Type: new Abstract: Joint OFDM-RIS optimization for 6G is a mixed-integer nonlinear programming (MINLP) problem covering sum-rate maximization, energy efficiency, max-min f

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

RoPoLL: Robust Panel of LLM Judges

DGX agent

arXiv:2606.30931v1 Announce Type: new Abstract: The LLM Jury, a Panel of LLM Evaluators (PoLL) reporting consensus scores, has become a practical alternative to single-judge LLM evaluation, yet its st

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action

DGX agent

arXiv:2606.31916v1 Announce Type: new Abstract: Theory of Mind (ToM) benchmarks for Large Language Models (LLMs) typically rely on passive question-answering formats, but the deployment of LLMs in inc

model-releasesarxiv-cs-cl
1 Jul 2026
Model Releases

When the Database Fails: Prompting LLM Dialogue Agents for Safe Recovery in Task-Oriented Dialogue

DGX agent

arXiv:2606.31307v1 Announce Type: new Abstract: Large language models used in task-oriented dialogue often produce fluent but unsafe responses when backend database calls fail, return empty results, o

model-releasesarxiv-cs-cl
1 Jul 2026
Model Releases

Bringing speed and strong cost performance to the market with Gemini Omni Flash and Nano Banana 2 Lite

DGX agent

Great creative happens when your tools move at the speed of your ideas. To help you create rich, reliable experiences while reducing regeneration time and costs, we’re adding two new models to Gemini

model-releasesgoogle-cloud-ai
30 Jun 2026
Model Releases

Can Fine-Tuning Erase Your Edits? On the Fragile Coexistence of Knowledge Editing and Adaptation

DGX agent

arXiv:2511.05852v4 Announce Type: replace-cross Abstract: Knowledge editing (KE) offers a lightweight alternative to retraining for updating large language models (LLMs). Meanwhile, fine-tuning remain

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Compressed Sensing for Capability Localization in Large Language Models

DGX agent

arXiv:2603.03335v2 Announce Type: replace Abstract: Large language models (LLMs) exhibit a wide range of capabilities, including mathematical reasoning, code generation, and linguistic behaviors. We s

model-releasesarxiv-cs-cl
30 Jun 2026
Model Releases

Defeat Devices in AI Systems

DGX agent

arXiv:2606.28863v1 Announce Type: cross Abstract: AI systems increasingly exhibit behavior that differs systematically between evaluation and deployment contexts. Alignment faking, sandbagging, benchm

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks

DGX agent

arXiv:2510.14207v3 Announce Type: replace Abstract: Large Language Model (LLM) agents are powering a growing share of interactive web applications, yet remain vulnerable to misuse and harm. Prior jail

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA

DGX agent

arXiv:2606.30220v1 Announce Type: new Abstract: High benchmark accuracy does not guarantee genuine use of visual evidence. We study this problem in traffic accident Video Question Answering (VideoQA),

model-releasesarxiv-cs-cv
30 Jun 2026
Model Releases

From Failure Taxonomy to Intervention: A Diagnostic Methodology for Industry-Scale AVLM in Video and Live-Streaming Platform Moderation

DGX agent

arXiv:2606.30059v1 Announce Type: new Abstract: Industry-scale video and live-streaming moderation imposes requirements that are difficult to satisfy with generic pretrained public models or external

model-releasesarxiv-cs-lg
30 Jun 2026
Model Releases

It Lied to a Doctor to Buy Poison Ingredients: Quantifying Real-World Misuse of Phone-use Agents

DGX agent

arXiv:2606.27944v1 Announce Type: cross Abstract: Phone-use Agents can execute complex tasks end to end across real mobile applications. By operating a real device on the user's behalf, they reach far

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

LLM-Guided Planning for Multi-hop Reasoning over Multimodal Nuclear Regulatory Documents

DGX agent

arXiv:2606.29399v1 Announce Type: new Abstract: Reviewing nuclear regulatory documents requires multi-hop reasoning across tens of thousands of pages, where judgments depend on evidence assembled acro

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Symbolic Mechanistic Data Attribution: Tracing Training Influence to Learned Behavioral Policies

DGX agent

arXiv:2606.29171v1 Announce Type: cross Abstract: While existing data attribution methods can identify which training examples build specific mechanistic circuits, they cannot explain how training dat

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Toward Secure and Reliable PDDL Formalization of Large Language Models with Planner-in-the-Loop Feedback

DGX agent

arXiv:2606.29700v1 Announce Type: new Abstract: Planning often requires symbolic specifications that are both executable and verifiable. For large language models deployed in autonomous or decision-su

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Whose Side Is Your Agent On? Multi-Party Principal Loyalty in LLM Agents

DGX agent

arXiv:2606.30383v1 Announce Type: new Abstract: A rapidly growing class of LLM agents is multi-party: the agent acts for a principal (who briefs it, sends follow-ups, and receives results) while also

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Aloe-Vision: Robust Vision-Language Models for Healthcare

DGX agent

arXiv:2606.27500v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) specialized in healthcare are emerging as a promising research direction due to their potential impact in clinica

model-releasesarxiv-cs-cl
29 Jun 2026
Industry

Could being labeled 'too dangerous' be good marketing for frontier AI firms? Hugging Face CEO Clem Delangue thinks so https://bloom.bg/4jP5R…

DGX agent

The Hugging Face CEO argues that being characterized as 'too dangerous' can serve as effective marketing for frontier AI companies, suggesting that such controversy may enhance their public profile an

industryclem-delangue--x
29 Jun 2026
Model Releases

FoggyTrust: Robust Federated Learning with Hierarchical Trust Networks

DGX agent

arXiv:2606.27622v1 Announce Type: new Abstract: Byzantine-robust federated learning seeks to protect distributed model training from malicious or corrupted clients without requiring access to their pr

model-releasesarxiv-cs-lg
29 Jun 2026
Model Releases

From Detection to Action: Using LLM Agents for Fault-Tolerant Control

DGX agent

arXiv:2606.28011v1 Announce Type: cross Abstract: We propose an agentic Large Language Model (LLM) framework for active Fault-Tolerant Control (FTC) that transforms fault detection outputs into constr

model-releasesarxiv-cs-lg
29 Jun 2026
Model Releases

iCost: A Novel Instance-Complexity-Based Cost-Sensitive Learning Framework

DGX agent

arXiv:2409.13007v3 Announce Type: replace-cross Abstract: Class imbalance poses a significant challenge in classification tasks, often causing standard learning algorithms to become biased toward the

model-releasesarxiv-cs-ai
29 Jun 2026
Industry

Open-source AI is booming, massively impactful for progress, competition, transparency & orders of magnitude less dangerous than closed-sour…

DGX agent

Clem Delangue argues that open-source AI development offers significant benefits for technological progress, market competition, and system transparency while presenting substantially lower risks comp

industryclem-delangue--x
29 Jun 2026
Model Releases

Is Gemini 3.5 Pro being export controlled? Because if not...

DGX agent

Ethan Mollick raises questions about whether Google's Gemini 3.5 Pro model should be subject to export controls, suggesting concerns about its capabilities and potential regulatory implications. The p

model-releasesethan-mollick--x
28 Jun 2026
Model Releases

VCs are now sharing screenshots in group chats of Claude discouraging investment in open-source AI infra startups and models. Obviously ther…

DGX agent

VCs are now sharing screenshots in group chats of Claude discouraging investment in open-source AI infra startups and models. Obviously there is an absolute EXPLOSION of pitches in inference companies

model-releasesclem-delangue--x
27 Jun 2026
Model Releases

Can Large Language Models Reliably Code Qualitative Humanitarian Data? A Benchmark Study Against Human Expert Adjudication

DGX agent

arXiv:2606.26541v1 Announce Type: new Abstract: Data from affected populations are crucial for informing humanitarian response, but their value depends on timely and consistent interpretation of nuanc

model-releasesarxiv-cs-lg
26 Jun 2026
Model Releases

ConvMemory v3: A Validity Context Layer for Conversational Memory via Target-Conditioned Relation Verification

DGX agent

arXiv:2606.26753v1 Announce Type: new Abstract: Conversational memory retrieval optimizes relevance, yet a retrieved memory can be relevant and simultaneously outdated: a later turn updates, corrects,

model-releasesarxiv-cs-cl
26 Jun 2026
Model Releases

Does AI Reviewer See the Full Picture? Attacking and Defending Multimodal Peer Review

DGX agent

arXiv:2606.12716v2 Announce Type: replace Abstract: The integration of Large Language Models (LLMs) and Multimodal LLMs (MLLMs) into scientific peer-review workflows introduces novel and significant r

model-releasesarxiv-cs-cl
26 Jun 2026
Local Ai

EndoUFM: Utilizing Foundation Models for Monocular depth estimation of endoscopic images

DGX agent

arXiv:2508.17916v2 Announce Type: replace Abstract: Depth estimation is a foundational component for 3D reconstruction in minimally invasive endoscopic surgeries. However, existing monocular depth est

local-aiarxiv-cs-cv
26 Jun 2026
← Previous
1…285286287288289…297
Next →