AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
Human
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
83,192 results
11 Aug 2026

Build it, Break it, Repeat: Benchmarking and improving LLM-manipulated disinformation detection in social media posts

Model ReleasesDGX agent

arXiv:2608.09510v1 Announce Type: cross Abstract: Detecting machine-generated disinformation on social media is increasingly difficult as large language models (LLMs) make it easier to generate and re

Building Agent Skills and testing them is hard, but it doesn't have to be. Listen to Arjun Patel demo Cultivar, an open source tool develope…

Model ReleasesDGX agent

Building Agent Skills and testing them is hard, but it doesn't have to be. Listen to Arjun Patel demo Cultivar, an open source tool developed at Pinecone to help benchmark agent skills in sandboxes. T

Business Arena: Benchmarking LLM Agents in a Realistic Marketplace

Agents
DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2608.08621v1 Announce Type: new Abstract: Running a business is a challenging form of intelligent work. Operators must infer opportunities from partial signals, commit capital under uncertainty,

Business Truth, not SQL Accuracy: A Rule-Gated 7B Analytics Agent Outperforms a Direct-Prompted 32B Baseline

AgentsDGX agent

arXiv:2608.09254v1 Announce Type: new Abstract: LLM analytics agents are evaluated on SQL syntax accuracy, but production failures look different: questions with two valid business definitions, questi

C^2A: Coupling Spatial Evidence with Clinical Priors via Co-occurrence Aware Class Attention for Multi-Label Chest X-Ray Classification

ResearchDGX agent

arXiv:2608.09774v1 Announce Type: new Abstract: Thoracic pathologies rarely occur in isolation, yet standard multi-label classifiers rely on shared global descriptors, discarding where findings lie an

CableDex: Cable Length Estimation on Industrial Reels Using a Handheld Device

ResearchDGX agent

arXiv:2608.09392v1 Announce Type: new Abstract: CableDex is a computer vision system that addresses the time-consuming and inaccurate manual measurement of cable length on industrial reels from a sing

CADEngBench: It Looks Like CAD, but Does It Work? Evaluating Parametric Design, Assembly Reasoning, and Physics Simulation

Model ReleasesDGX agent

arXiv:2608.09296v1 Announce Type: new Abstract: A CAD model is not engineering-grade merely because it looks correct. It must satisfy design requirements, respond predictably to parameter changes, sup

Calib3R: Hand-Eye Calibration and 3D Metric-Scaled Scene Reconstruction with 3D Foundation Models

ResearchDGX agent

arXiv:2509.08813v2 Announce Type: replace Abstract: Robots often rely on RGB images for tasks like manipulation. However, reliable interaction typically requires a 3D scene representation that is metr

Calling the Bluff: Detecting Ever-Shifting Harmful Chat Dialogue via Ordered Reasoning Chain Regularization

ResearchDGX agent

arXiv:2608.08451v1 Announce Type: cross Abstract: Harmful chat dialogues are ever-shifting through type-shifting and lexical evasion, yet we find they share invariant principles, i.e., an Ordered Reas

Can Coding Agents Solve Repository-Level Issues with Rendered Code? An Exploratory Study of Visual Representations

AgentsDGX agent

arXiv:2608.09268v1 Announce Type: cross Abstract: Visual modality has recently been explored as a way to compress textual tokens, including rendering code as images for static code understanding. We s

Can Gemma and Qwen models catch hallucinations by looking at their own logprobs?

Model ReleasesDGX agent

Hi! I'm really obsessed with LLM hallucinations for the last 6 days 😭 I started by designing system prompts to attack hallucinations but failed, obviously. Now I tried reading logprobs and... I think

Can Graph Learning Learn Circuits?

Model ReleasesDGX agent

arXiv:2608.08536v1 Announce Type: new Abstract: Circuit localization is a mechanistic interpretability task whose goal is to identify a sparse subgraph of a transformer's computation graph sufficient

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives

Model ReleasesDGX agent

arXiv:2608.08160v1 Announce Type: cross Abstract: The rapid advancement of Large Language Models (LLMs) is revolutionizing AI for Games by enabling open-ended and fluid interactive storytelling. Howev

Can Open-Weight Models Compete on Financial Text Comprehension?

Model ReleasesDGX agent

arXiv:2608.08634v1 Announce Type: new Abstract: Open-weight language models from Chinese AI labs caught up on benchmarks relative to proprietary frontier models in recent months. Yet their reliability

Can We Optimize the Performance-Carbon Emission Break-Even Point?: The Quest for Greener LLMs

Model ReleasesDGX agent

arXiv:2608.08744v1 Announce Type: cross Abstract: The carbon footprint of any deployed Large Language Model (LLM) accumulates during inference, where repeated use of the model substantially exceeds th

Can Webcam Gaze Constrain Mesa-Objectives in Driving Models? An Instrument Precision Analysis

AgentsDGX agent

arXiv:2608.08947v1 Announce Type: new Abstract: Current hazard detection systems in autonomous driving may develop mesa objectives, learned internal goals that achieve high training performance throug

CAP: A Scalable Benchmark for Evaluating Cross-Site Browser Agents with Complex Actions and Perception

Model ReleasesDGX agent

arXiv:2608.08392v1 Announce Type: new Abstract: Large language models are increasingly deployed as autonomous agents that interact with the web through browsers. While recent progress has been driven

Capability Is Not Propensity: Measuring Pressure-Robust Cooperative Behavior in Civic LLM Agents

ResearchDGX agent

arXiv:2608.09485v1 Announce Type: new Abstract: Cooperative capabilities in language models are dual-use. The same social reasoning that supports civic deliberation can also enable strategic omission,

CARD: Controlled Agentic Reddit Discussions for Credit Card Simulation

AgentsDGX agent

arXiv:2608.09790v1 Announce Type: new Abstract: Online credit card discussions provide a natural setting for studying how consumers communicate about financial products. Simulating these discussions r

Carnot: Interpretable, Interactive, and Optimized Execution of Deep Research Queries

ApplicationsDGX agent

arXiv:2608.09532v1 Announce Type: cross Abstract: Enterprises increasingly seek to query data lakes using natural language via AI-driven tools like semantic operators or deep research agents. However,

Catastrophic Forgetting in Continual Reinforcement Learning

ResearchDGX agent

arXiv:2608.08673v1 Announce Type: new Abstract: This work explores the relationship between task similarity and catastrophic forgetting in reinforcement learning. Catastrophic forgetting, the phenomen

Causal Falsification of Digital Twins

SafetyDGX agent

arXiv:2301.07210v5 Announce Type: replace-cross Abstract: Digital twins are simulation-based models designed to predict how a real-world process will evolve in response to interventions. This modellin

Causal State-Space Model for Causal Inference: Estimating Longitudinal Individual Treatment Effects

ResearchDGX agent

arXiv:2608.08288v1 Announce Type: new Abstract: Estimating counterfactual outcomes over time from longitudinal observational data is central to clinical decision support. Existing methods rely on doma

CausalNav: Reliability-Certified Causal World Models for Control under Physical-Parameter Shift

Model ReleasesDGX agent

arXiv:2608.07809v1 Announce Type: new Abstract: A world model is only useful for physical AI if it changes what the agent does, and only safe if it declines to do so when it is wrong. We study both ha

CDGC-Net: 3D Medical Image Segmentation with Cooperative Dual-Scale Self-Attention and Grouped Channel Modeling

SafetyDGX agent

arXiv:2608.08575v1 Announce Type: cross Abstract: Accurate 3D medical image segmentation requires the integration of long-range anatomical context with fine boundary detail. Existing methods often mod

CEAA: A Cognitive Embodied Agents Architecture for Interactive Computing Systems

AgentsDGX agent

arXiv:2608.09848v1 Announce Type: new Abstract: The development of embodied Intelligent Virtual Agents (IVAs) that have cognitive capabilities in real-time interactive virtual environments remains a c

CFD-Guided Detection of Concept Drift in Multimodal Physiologic Signals

ResearchDGX agent

arXiv:2608.07759v1 Announce Type: cross Abstract: Cardiovascular AI models can classify clean elec- trocardiogram (ECG) signals, but real wearable signals change because of motion, breathing, posture,

CGRL: Causal-Guided Representation Learning for Node-Level Out-of-Distribution Generalization

Model ReleasesDGX agent

arXiv:2603.24304v2 Announce Type: replace-cross Abstract: Graph Neural Networks (GNNs) deliver strong performance on graph tasks, but their accuracy drops significantly under out-of-distribution (OOD)

ChatGPT and Gemini both just passed 1 billion users

Model ReleasesDGX agent

For the 14th time, a Google product has hit 1 billion users. Google CEO Sundar Pichai posted on X that a billion people are using Gemini every month, and that Gemini is Google's fastest-growing produc

ChronoState: Hidden Elapsed-Time Conditioning for Temporal-State Action Selection in Frozen-Backbone Language Models

Model ReleasesDGX agent

arXiv:2608.09124v1 Announce Type: new Abstract: Temporal decisions in language-model systems often depend on both symbolic task state and elapsed wall-clock time, such as cache expiration, job complet

CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment

Model ReleasesDGX agent

arXiv:2608.09164v1 Announce Type: new Abstract: Aligning large language models (LLMs) with human privacy preferences requires capturing individuals' disclosure boundaries beyond general privacy norms.

CIFA: Contextual-Intersectional Fairness Auditing for Hidden Subgroup Discovery in Face Analysis

SafetyDGX agent

arXiv:2608.09669v1 Announce Type: new Abstract: Fairness evaluation in computer vision commonly relies on aggregate accuracy and demographic subgroup analysis. However, visual models are also sensitiv

Circuit Fine-Tuning for Compute-Efficient Transformer Adaptation

Model ReleasesDGX agent

arXiv:2608.08336v1 Announce Type: new Abstract: Parameter-Efficient Fine-Tuning (PEFT) has become the de facto standard for adapting Vision Transformers (ViTs) to downstream tasks. While parameter cou

CircuitReason-1k: Benchmarking Long-Horizon Visual-to-Symbolic Reasoning inElectrical Circuits

Model ReleasesDGX agent

arXiv:2608.09374v1 Announce Type: new Abstract: Electrical circuit analysis requires more than recognizing components in an image. A solver must ground symbols and labels, recover latent topology, sel

City Sentinel: A Unified AI-Based Smart Surveillance Framework for Real-Time Multi-Threat Detection Using Deep Learning

SafetyDGX agent

arXiv:2608.08887v1 Announce Type: new Abstract: Rapid urbanization has increased the need for surveillance systems that can monitor multiple public safety risks at the same time. Traditional systems o

CLAM: Causal Spatial Disaggregation to Infer Local Effects From Coarse Data

SafetyDGX agent

arXiv:2608.08064v1 Announce Type: new Abstract: Learning fine-grained spatial patterns from coarse-resolution data is challenging, especially in causal settings where high-resolution effects must be i

Classical SU(2) Models Match or Exceed Shallow Variational Quantum Circuits on Vision Benchmarks

Local AiDGX agent

arXiv:2608.07822v1 Announce Type: cross Abstract: Quaternion-valued neural networks and variational quantum circuits (VQCs) both derive local transformations from SU(2) geometry, yet their performance

Claude will apply invisible watermarks to AI text and images

Model ReleasesDGX agent

Anthropic has pledged to start marking Claude-generated text and images with machine-readable data, in an effort to comply with European rules for AI transparency. 'Generated text will carry embedded

Claude's watermark probably doesn't work how you think. As the CTO of GPTZero, I'll explain how Anthropic, Google and OpenAI are building te…

Model ReleasesDGX agent

Claude's watermark probably doesn't work how you think. As the CTO of GPTZero, I'll explain how Anthropic, Google and OpenAI are building text watermarking in this brief explainer and whether it can b

CLEAR: A Semantic-Geometric Terrain Abstraction for Large-Scale Unstructured Environments

AgentsDGX agent

arXiv:2601.13361v2 Announce Type: replace Abstract: Long-horizon navigation in unstructured environments demands terrain abstractions that scale to tens of square kilometers while preserving semantic

CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR

Model ReleasesDGX agent

arXiv:2608.07796v1 Announce Type: new Abstract: Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment requires agents to conduct defensible investiga

Closing the loop in learning with missing data

Model ReleasesDGX agent

arXiv:2608.09030v1 Announce Type: cross Abstract: What should a machine learning model learn when data is missing during training? We look at the learning process from a dynamical systems perspective,

Clustered Attractor Manifolds and Dynamical Condensation in Self-Attention

ResearchDGX agent

arXiv:2608.08922v1 Announce Type: cross Abstract: Transformer layers generate state-dependent interaction networks: token representations determine the attention matrix, which in turn updates the repr

CMU-Drive and V2V-VLA: Cooperative Multi-agent Unified Driving with Reasoning Benchmark and Vehicle-to-Vehicle Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2608.07621v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have recently achieved impressive performance for end-to-end autonomous driving, yet existing approaches are primari

Coarse-to-Fine Registration of Jawbone CT and Intraoral Scan Data Using GeDi and ICP with Pseudo-IOS Ground Truth

SafetyDGX agent

arXiv:2608.07564v1 Announce Type: cross Abstract: In digital dentistry and oral surgery, the registration of jawbone CT and intraoral scanner (IOS) data is essential for integrating internal bone stru

CoCoNav: Conformal Control for Safe Robot Navigation in Crowds

ResearchDGX agent

arXiv:2608.07751v1 Announce Type: cross Abstract: Safe and efficient robot navigation in crowds requires anticipating pedestrian motion despite uncertain and potentially shifting prediction errors. Ex

CodecArena: Codec Quality Assessment via Visual Reinforcement Learning

Model ReleasesDGX agent

arXiv:2608.09139v1 Announce Type: new Abstract: Video coding is advancing into the low and ultra-low bitrate regime, driven by end-to-end codecs that replace the hand-crafted pipeline with jointly opt

CODS: Iterative Bellman-Residual Data Selection for Reusable Offline Reinforcement Learning

ResearchDGX agent

arXiv:2608.07719v1 Announce Type: new Abstract: Offline reinforcement learning repeatedly trains policies from a fixed transition pool, making redundant data costly across seeds and hyperparameters, w

Cognitive Energy Modeling for Neuroadaptive Human-Machine Systems using EEG and WGAN-GP

ResearchDGX agent

arXiv:2604.01653v2 Announce Type: replace Abstract: Electroencephalography (EEG) provides a non-invasive insight into the brain's cognitive and emotional dynamics. However, modeling how these states e

CoInS-Net: A Continuous Position-Aware Network for Joint Medical Image Interpolation and Segmentation

ResearchDGX agent

arXiv:2608.09391v1 Announce Type: new Abstract: Accurate medical image interpolation and anatomical structure segmentation are fundamental for computer-aided diagnosis and treatment planning. Anisotro

ColluSkill: Adversarial Cross-Skill Composition for Evading Agent Skill Scanners

Local AiDGX agent

arXiv:2608.09732v1 Announce Type: cross Abstract: Agent skills are emerging as an important attack surface in LLM-based agent systems. Through an empirical study of existing skill scanners, we find th

ComboShoppingBench: Evaluating LLM Agents for Budget-Constrained Basket Shopping with Coupons

Model ReleasesDGX agent

arXiv:2608.09282v1 Announce Type: new Abstract: Real-world shopping often requires constructing a basket of complementary items rather than retrieving a single product. Such combo-shopping tasks arise

COMEX: A Composition-Grounded Benchmark and Learning Framework for Explainable Aesthetic Image Cropping

Model ReleasesDGX agent

arXiv:2608.07570v1 Announce Type: cross Abstract: Explainable aesthetic image cropping requires not only localizing a visually pleasing crop but also explaining why it is preferred. Existing crop-and-

CommitKV: Lifecycle-Aware KV Cache Compression via Commit Transitions for Multi-Turn Agents

AgentsDGX agent

arXiv:2608.07855v1 Announce Type: new Abstract: Multi-turn Reasoning-and-Acting (ReAct) agents accumulate growing trajectories of reasoning, tool calls, and observations. Their key-value (KV) caches g

Commitment Before Realization: When Classifier-Free Guidance Becomes Unnecessary in Masked Diffusion Language Models

ResearchDGX agent

arXiv:2608.08082v1 Announce Type: new Abstract: Classifier-free guidance (CFG) is usually kept on throughout masked diffusion language model decoding, although its benefit varies across prompts and ov

Communication-efficient distributed hazard difference estimation for heterogeneous multi-site survival data

ResearchDGX agent

arXiv:2601.14609v2 Announce Type: replace-cross Abstract: Multi-site collaboration can power survival models that no single hospital could fit alone, but privacy rules and protected computing environm

Communication Outage-Resistant UUV State Estimation: A Variational History Distillation Approach

ResearchDGX agent

arXiv:2603.29512v3 Announce Type: replace Abstract: The reliable operation of Unmanned Underwater Vehicle (UUV) clusters is highly dependent on continuous acoustic communication. However, this communi

Comparing British and American Audio Description of Movies

ResearchDGX agent

arXiv:2608.09792v1 Announce Type: new Abstract: Narrating the visual component of movies is known as audio description. It is a narrative technique designed to enable blind and visually impaired indiv

Compiling and Benchmarking Task-State Horizons for Embodied Agents

Model ReleasesDGX agent

arXiv:2608.08036v1 Announce Type: new Abstract: Frontier agentic models are increasingly deployed as high-level planners for long-horizon embodied tasks. Existing robotic benchmarks have advanced long

Complete, Scalable, and Robust Prioritized Planning for Multi-Robot Ordered Storage and Retrieval at Maximum Capacity

AgentsDGX agent

arXiv:2608.07734v1 Announce Type: cross Abstract: Automated warehouses face a fundamental trade-off between maximizing storage density and achieving high retrieval throughput. While puzzle-based stora

← Previous
1…1415161718…1387
Next →