AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
83,164 results
11 Aug 2026

BDH-CQ: In-Context Learning with Recurrent Latent Reasoning

Model ReleasesDGX agent

arXiv:2608.09888v1 Announce Type: cross Abstract: We introduce BDH-CQ, a reasoning model that combines in-context learning with recurrent latent reasoning. Inputs presented at inference time continuou

Benchmarking In-context Experiential Learning Through Repeated Product Recommendations

Model ReleasesDGX agent

arXiv:2511.22130v2 Announce Type: replace Abstract: To navigate ever-shifting real-world environments, agents must grapple with incomplete knowledge and adapt their strategies through experience. Howe

Benchmarking the Robustness of Agentic Systems to Adversarially-Induced Harms

Model ReleasesDGX agent

arXiv:2508.16481v3 Announce Type: replace Abstract: Ensuring the safe use of agentic systems requires a thorough understanding of the range of malicious behaviors these systems may exhibit. In this pa

Content type
AllBlogX PostPaperYouTubeRedditGitHub

Beyond Aggregate Calibration: Decomposing Income-Conditional Recall Disparities in Automated Credit Default Prediction

SafetyDGX agent

arXiv:2608.08202v1 Announce Type: new Abstract: Data-centric curation pipelines frequently rely on model confidence scores to flag and filter noisy or mislabeled training instances. Evaluating this fi

Beyond Binary: Continuous State Optimization with Graph-Structured Objectives

SafetyDGX agent

arXiv:2608.09366v1 Announce Type: new Abstract: Large-scale learning systems often face the challenge of balancing multiple, potentially competing objectives, such as fairness, accuracy, and latency.

Beyond cognacy

SafetyDGX agent

arXiv:2507.03005v3 Announce Type: replace Abstract: Computational phylogenetics has become an established tool in historical linguistics, with many language families now analyzed using likelihood-base

Beyond Direct Identifiers: Probabilistic Privacy Risk Estimation for Privacy-Conscious LLM Query Delegation

Model ReleasesDGX agent

arXiv:2608.09140v1 Announce Type: cross Abstract: Recent work on protecting privacy during user-LLM interactions often focuses on direct, explicit identifiers: the personally-identifiable information

Beyond Global Editing: Per-Instance Disentangled Subspaces for Training-Free Hallucination Mitigation in LVLMs

ResearchDGX agent

arXiv:2608.09344v1 Announce Type: new Abstract: Recent advances in large vision-language models (LVLMs) have enabled powerful multimodal reasoning by integrating visual encoders with large language mo

Beyond Hazard Resemblance: Contrastive Event Adjudication for Training-Free Video Anomaly Detection

Local AiDGX agent

arXiv:2608.09908v1 Announce Type: new Abstract: Video anomaly detection (VAD) aims to identify and temporally localize abnormal events in videos. Supervised methods learn anomaly decision boundaries f

Beyond 'I Can't Help With That': How Child Safety Experts Evaluate AI Chatbot Safety

SafetyDGX agent

arXiv:2608.07902v1 Announce Type: cross Abstract: Youth increasingly turn to AI chatbots for social and emotional support, raising concerns about how these systems respond, especially in high-stakes s

Beyond Isotropic Assumptions: Continuity-Constrained Segmentation and GPU Morphometry for Nanoscale GBM Analysis

HardwareDGX agent

arXiv:2608.07575v1 Announce Type: new Abstract: Confocal microscopy of optically cleared and swelled tissue resolves complex biological structures in 3D, but such acquisitions are highly anisotropic:

Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions

Model ReleasesDGX agent

arXiv:2608.09930v1 Announce Type: cross Abstract: Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expecte

Beyond Pixels: Benchmarking and Reward-Based Assessing Framework for Visual Spatial Aesthetics

Model ReleasesDGX agent

arXiv:2512.05098v2 Announce Type: replace-cross Abstract: In recent years, Image Quality Assessment (IQA) for AI-generated images (AIGI) has advanced rapidly; however, existing methods primarily targe

Beyond Pixels: Exploring DOM Downsampling for LLM-Based Web Agents

Model ReleasesDGX agent

arXiv:2508.04412v3 Announce Type: replace Abstract: The advent of large language models (LLMs) has sparked an evolution of autonomous web browsing agents: given a web browsing task and serialised user

Beyond Routing: Decoupling Expert Dispatch and Aggregation in Sparse Mixture-of-Experts

Model ReleasesDGX agent

arXiv:2608.08853v1 Announce Type: new Abstract: Sparse Mixture-of-Experts (MoE) routers commonly use the same scores both to select experts and to weight their already-computed outputs. We study wheth

Beyond Solvability: Task Learnability as a Static Prior for LLM RL Post-Training

SafetyDGX agent

arXiv:2608.09217v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a central post-training paradigm for eliciting reasoning capabilities in large language models, yet uniform tas

Beyond Static Models: An Evolving Framework for Continual Learning in Large Language Models across Training Stages

Model ReleasesDGX agent

arXiv:2603.12658v2 Announce Type: replace-cross Abstract: Continual learning (CL) has emerged as a pivotal paradigm to enable large language models (LLMs) to dynamically adapt to evolving knowledge an

Beyond Tables: Doc2DB-Bench for Relationally Faithful Document-to-Database Construction

Model ReleasesDGX agent

arXiv:2608.08459v1 Announce Type: cross Abstract: Practical AI systems increasingly need to turn long, heterogeneous documents into queryable relational databases, not isolated spreadsheets. In domain

Beyond the Capability Boundary: Zeroth-Order Optimization for Self-Evolving LLM Agents

Model ReleasesDGX agent

arXiv:2608.09292v1 Announce Type: cross Abstract: Self-evolving methods improve the capabilities of LLM agents by sampling trajectories from the underlying LLMs and learning from these trajectories. H

Beyond the Node: Clade-level Selection for Efficient MCTS in Automatic Heuristic Design

TutorialsDGX agent

arXiv:2602.00549v2 Announce Type: replace Abstract: While Monte Carlo Tree Search (MCTS) shows promise in Large Language Model (LLM) based Automatic Heuristic Design (AHD), it suffers from a critical

Beyond the Plane: Coupling Planar Vehicle Dynamics with Three-Dimensional Road Geometry

Local AiDGX agent

arXiv:2608.09402v1 Announce Type: new Abstract: Simulation is crucial for developing and testing autonomous driving systems. In particular, the development of localization and control algorithms relie

Beyond Uniform Restoration: Empowering All-in-One Restoration with Pixel-Level Multimodal Guidance

Local AiDGX agent

arXiv:2608.09482v1 Announce Type: cross Abstract: All-in-one image restoration is a unified low-level vision task that aims to effectively recover high-quality images from inputs degraded by various t

BibTeX Citation Errors in Scientific Publishing Agents: Evaluation and Mitigation

Model ReleasesDGX agent

arXiv:2604.03159v2 Announce Type: replace-cross Abstract: Large language models with web search are increasingly used in scientific publishing agents, yet they produce BibTeX entries with pervasive fi

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents

Model ReleasesDGX agent

arXiv:2608.09555v1 Announce Type: new Abstract: External natural-language skills provide large language model (LLM) agents with reusable and editable guidance for solving complex tasks. Yet their effe

Big Tech's AI boom echoes the 1870s railroad buildout, and Nvidia shifting risk to institutional capital may expose investors if AI revenues fail to materialize (Ben Thompson/Stratechery)

HardwareDGX agent

Ben Thompson / Stratechery: Big Tech's AI boom echoes the 1870s railroad buildout, and Nvidia shifting risk to institutional capital may expose investors if AI revenues fail to materialize — On Januar

Biologically Informed Representation Learning for Robust Cross-Center Generalization of MALDI-TOF Mass Spectrometry

Model ReleasesDGX agent

arXiv:2608.08182v1 Announce Type: cross Abstract: Machine learning models for MALDI-TOF mass spectrometry have shown considerable promise for clinical microbiology tasks such as microbial identificati

Blumira launches Hearth, an AI command center that spans rival security tools

Model ReleasesDGX agent

Security operations platform startup Blumira Inc. today launched Hearth, a vendor-agnostic artificial intelligence command center that lets security teams investigate and act across their existing too

BMDS-Net:Deployment-aware multi-modal brain tumor segmentation with adaptive fusion,decoder regularization,and Bayesian calibration

ResearchDGX agent

arXiv:2601.17504v2 Announce Type: replace Abstract: Multi-modal MRI enables detailed brain tumor sub-region segmentation, but clinical deployment remains affected by missing sequences,boundary errors,

Bootstrapping Vision-Language Model for Hysteroscopic Surgical Scene Segmentation

Local AiDGX agent

arXiv:2608.09302v1 Announce Type: new Abstract: Hysteroscopic surgical scene segmentation plays a pivotal role in understanding the hysteroscopic intraoperative environment as well as computer-assiste

Bounding Hallucinations: Merlin-Arthur Protocols for Mutual-Information Bounds in Language Models

AgentsDGX agent

arXiv:2512.11614v3 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) relies on retrieved context to guide large language models (LLM), yet treats the retrieval as a heuristic

BRACE: Taming Sharp Irregularities via Barycentric Rational Forecasting for Fast Diffusion Transformers Inference

ResearchDGX agent

arXiv:2608.07572v1 Announce Type: cross Abstract: Diffusion Transformers (DiTs) have demonstrated exceptional performance in high-fidelity image and video generation. To alleviate their massive comput

Branch2Skill: Efficient Skill Evolution Through Reasoning Trees

AgentsDGX agent

arXiv:2608.08677v1 Announce Type: new Abstract: Skill evolution improves agent skills through feedback over time, with failed trajectories often providing informative signals by revealing incomplete o

Bridging Object Detection and Segmentation with Polygon Detection Transformers

AgentsDGX agent

arXiv:2603.09245v2 Announce Type: replace Abstract: Box detection and mask segmentation are two dominant paradigms for foreground representation: boxes are efficient but too coarse for object shapes,

Bridging the Evaluation Gap: Standardized Benchmarks for Multi-Objective Search

Model ReleasesDGX agent

arXiv:2603.24084v2 Announce Type: replace Abstract: Empirical evaluation in multi-objective search (MOS) has historically suffered from fragmentation, relying on heterogeneous problem instances with i

Bridging the Gap Between Semantics and Reconstruction:Unifying Sign Language Translation and Production

ApplicationsDGX agent

arXiv:2608.09045v1 Announce Type: cross Abstract: Recent advances in sign language (SL) research have shown a trend toward unifying multiple sign language understanding (SLU) subtasks, such as isolate

Bright-Channel Retinex Enhancement with a Conditional Overdispered-Noise Analysis

Local AiDGX agent

arXiv:2608.09137v1 Announce Type: new Abstract: I present a training-free low-light enhancement method that combines local bright-channel illumination estimation, Retinex division, and edge-preserving

Brownian Kernel Ladders

ResearchDGX agent

arXiv:2606.15812v2 Announce Type: replace Abstract: We introduce Brownian kernel ladders (BKLs), a recursive hierarchy of integral reproducing kernel Hilbert spaces built from linear functionals by re

BRUCE: Benchmarking Robustness Under Corruption Escalation for Scientific Vision-Language Reasoning

ApplicationsDGX agent

arXiv:2608.07742v1 Announce Type: new Abstract: Visual-language models (VLMs) frequently struggle with robustness issues in real-world situations due to low- or varying-quality input images. In this p

Build it, Break it, Repeat: Benchmarking and improving LLM-manipulated disinformation detection in social media posts

Model ReleasesDGX agent

arXiv:2608.09510v1 Announce Type: cross Abstract: Detecting machine-generated disinformation on social media is increasingly difficult as large language models (LLMs) make it easier to generate and re

Building Agent Skills and testing them is hard, but it doesn't have to be. Listen to Arjun Patel demo Cultivar, an open source tool develope…

Model ReleasesDGX agent

Building Agent Skills and testing them is hard, but it doesn't have to be. Listen to Arjun Patel demo Cultivar, an open source tool developed at Pinecone to help benchmark agent skills in sandboxes. T

Business Arena: Benchmarking LLM Agents in a Realistic Marketplace

AgentsDGX agent

arXiv:2608.08621v1 Announce Type: new Abstract: Running a business is a challenging form of intelligent work. Operators must infer opportunities from partial signals, commit capital under uncertainty,

Business Truth, not SQL Accuracy: A Rule-Gated 7B Analytics Agent Outperforms a Direct-Prompted 32B Baseline

AgentsDGX agent

arXiv:2608.09254v1 Announce Type: new Abstract: LLM analytics agents are evaluated on SQL syntax accuracy, but production failures look different: questions with two valid business definitions, questi

C^2A: Coupling Spatial Evidence with Clinical Priors via Co-occurrence Aware Class Attention for Multi-Label Chest X-Ray Classification

ResearchDGX agent

arXiv:2608.09774v1 Announce Type: new Abstract: Thoracic pathologies rarely occur in isolation, yet standard multi-label classifiers rely on shared global descriptors, discarding where findings lie an

CableDex: Cable Length Estimation on Industrial Reels Using a Handheld Device

ResearchDGX agent

arXiv:2608.09392v1 Announce Type: new Abstract: CableDex is a computer vision system that addresses the time-consuming and inaccurate manual measurement of cable length on industrial reels from a sing

CADEngBench: It Looks Like CAD, but Does It Work? Evaluating Parametric Design, Assembly Reasoning, and Physics Simulation

Model ReleasesDGX agent

arXiv:2608.09296v1 Announce Type: new Abstract: A CAD model is not engineering-grade merely because it looks correct. It must satisfy design requirements, respond predictably to parameter changes, sup

Calib3R: Hand-Eye Calibration and 3D Metric-Scaled Scene Reconstruction with 3D Foundation Models

ResearchDGX agent

arXiv:2509.08813v2 Announce Type: replace Abstract: Robots often rely on RGB images for tasks like manipulation. However, reliable interaction typically requires a 3D scene representation that is metr

Calling the Bluff: Detecting Ever-Shifting Harmful Chat Dialogue via Ordered Reasoning Chain Regularization

ResearchDGX agent

arXiv:2608.08451v1 Announce Type: cross Abstract: Harmful chat dialogues are ever-shifting through type-shifting and lexical evasion, yet we find they share invariant principles, i.e., an Ordered Reas

Can Coding Agents Solve Repository-Level Issues with Rendered Code? An Exploratory Study of Visual Representations

AgentsDGX agent

arXiv:2608.09268v1 Announce Type: cross Abstract: Visual modality has recently been explored as a way to compress textual tokens, including rendering code as images for static code understanding. We s

Can Gemma and Qwen models catch hallucinations by looking at their own logprobs?

Model ReleasesDGX agent

Hi! I'm really obsessed with LLM hallucinations for the last 6 days 😭 I started by designing system prompts to attack hallucinations but failed, obviously. Now I tried reading logprobs and... I think

Can Graph Learning Learn Circuits?

Model ReleasesDGX agent

arXiv:2608.08536v1 Announce Type: new Abstract: Circuit localization is a mechanistic interpretability task whose goal is to identify a sparse subgraph of a transformer's computation graph sufficient

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives

Model ReleasesDGX agent

arXiv:2608.08160v1 Announce Type: cross Abstract: The rapid advancement of Large Language Models (LLMs) is revolutionizing AI for Games by enabling open-ended and fluid interactive storytelling. Howev

Can Open-Weight Models Compete on Financial Text Comprehension?

Model ReleasesDGX agent

arXiv:2608.08634v1 Announce Type: new Abstract: Open-weight language models from Chinese AI labs caught up on benchmarks relative to proprietary frontier models in recent months. Yet their reliability

Can We Optimize the Performance-Carbon Emission Break-Even Point?: The Quest for Greener LLMs

Model ReleasesDGX agent

arXiv:2608.08744v1 Announce Type: cross Abstract: The carbon footprint of any deployed Large Language Model (LLM) accumulates during inference, where repeated use of the model substantially exceeds th

Can Webcam Gaze Constrain Mesa-Objectives in Driving Models? An Instrument Precision Analysis

AgentsDGX agent

arXiv:2608.08947v1 Announce Type: new Abstract: Current hazard detection systems in autonomous driving may develop mesa objectives, learned internal goals that achieve high training performance throug

CAP: A Scalable Benchmark for Evaluating Cross-Site Browser Agents with Complex Actions and Perception

Model ReleasesDGX agent

arXiv:2608.08392v1 Announce Type: new Abstract: Large language models are increasingly deployed as autonomous agents that interact with the web through browsers. While recent progress has been driven

Capability Is Not Propensity: Measuring Pressure-Robust Cooperative Behavior in Civic LLM Agents

ResearchDGX agent

arXiv:2608.09485v1 Announce Type: new Abstract: Cooperative capabilities in language models are dual-use. The same social reasoning that supports civic deliberation can also enable strategic omission,

CARD: Controlled Agentic Reddit Discussions for Credit Card Simulation

AgentsDGX agent

arXiv:2608.09790v1 Announce Type: new Abstract: Online credit card discussions provide a natural setting for studying how consumers communicate about financial products. Simulating these discussions r

Carnot: Interpretable, Interactive, and Optimized Execution of Deep Research Queries

ApplicationsDGX agent

arXiv:2608.09532v1 Announce Type: cross Abstract: Enterprises increasingly seek to query data lakes using natural language via AI-driven tools like semantic operators or deep research agents. However,

Catastrophic Forgetting in Continual Reinforcement Learning

ResearchDGX agent

arXiv:2608.08673v1 Announce Type: new Abstract: This work explores the relationship between task similarity and catastrophic forgetting in reinforcement learning. Catastrophic forgetting, the phenomen

Causal Falsification of Digital Twins

SafetyDGX agent

arXiv:2301.07210v5 Announce Type: replace-cross Abstract: Digital twins are simulation-based models designed to predict how a real-world process will evolve in response to interventions. This modellin

← Previous
1…1314151617…1387
Next →