AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,965
  • Agents7,446
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,131
  • Local Ai4,857
  • Model Releases23,360
  • Research19,834
  • Safety13,174
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,965
  • Agents7,446
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,131
  • Local Ai4,857
  • Model Releases23,360
  • Research19,834
  • Safety13,174
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
HumanDGX agent

Content type
86,965Total entries
1Added by human
86,964Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
51,106 results
Model Releases

ACON: Optimizing Context Compression for Long-horizon LLM Agents

DGX agent

arXiv:2510.00615v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed as agents in dynamic real-world environments, where success depends on maintaining precise re

model-releasesarxiv-cs-ai
2 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

An Open-Source Benchmark and Baseline for Multi-temporal Referring Segmentation

DGX agent

arXiv:2606.00987v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) have shown strong visual understanding and language-guided grounding abilities, yet their capacity for multi-temp

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

APEX-SQL: Talking to the data via Agentic Exploration for Text-to-SQL

DGX agent

arXiv:2602.16720v2 Announce Type: replace-cross Abstract: Text-to-SQL systems powered by Large Language Models have excelled on academic benchmarks but struggle in complex enterprise environments. The

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

ATLAS: Agentic Test-time Learning-to-Allocate Scaling

DGX agent

arXiv:2606.01667v1 Announce Type: new Abstract: Test-time scaling has become a major way to improve large language model reasoning, but its orchestration has remained designer-engineered: a fixed samp

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

Benchmarking Multimodal LLMs on Code Generation for Complex Interactive Webpages

DGX agent

arXiv:2606.00154v1 Announce Type: cross Abstract: Recent advancements in multimodal large language models (MLLMs) have achieved remarkable progress in multimodal reasoning and code generation, catalyz

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Can LLMs Reason Structurally? Benchmarking via the Lens of Data Structures

DGX agent

arXiv:2505.24069v4 Announce Type: replace-cross Abstract: Large language models (LLMs) are deployed on increasingly complex tasks that require multi-step decision-making. Understanding their algorithm

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

CityTrajBench: A Unified Benchmark for City-Scale Vehicle Trajectory Generation

DGX agent

arXiv:2606.02287v1 Announce Type: cross Abstract: Urban trajectory generation is a fundamental task for transportation simulation, urban planning, and mobility analytics. However, systematic compariso

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Connecting the Dots: Benchmarking Reflective Memory in Long-Horizon Dialogue

DGX agent

arXiv:2606.01223v1 Announce Type: cross Abstract: Despite substantial progress in long-context modeling, existing benchmarks remain confined to factual memory for explicit recall, failing to measure t

model-releasesarxiv-cs-ai
2 Jun 2026
Safety

Consistency Training while Mitigating Obfuscation via Rate Matching

DGX agent

arXiv:2606.02211v1 Announce Type: cross Abstract: Large language models are often influenced by extraneous input features, such as cues revealing a user's preferred answer. Consistency training reduce

safetyarxiv-cs-ai
2 Jun 2026
Model Releases

Consistent and Distinctive: LLM Benchmark Efficiency via Maximum Independent Set Prompt Selection on Similarity Graphs

DGX agent

arXiv:2606.01400v1 Announce Type: cross Abstract: Evaluating large language models (LLMs) across comprehensive benchmarks is expensive and time-consuming. We propose a graph-based prompt selection fra

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Decision-Focused On-Policy Learning for Contextual Linear Optimization with Partial Feedback

DGX agent

arXiv:2606.01081v1 Announce Type: new Abstract: Decision-focused learning (DFL) trains predictive models by optimizing downstream decision quality rather than standalone prediction accuracy. For conte

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

Density-Aware Translation of Spurious Correlations in Zero-Shot VLMs

DGX agent

arXiv:2606.01710v1 Announce Type: new Abstract: Vision-Language models (VLMs), such as CLIP, achieve powerful zero-shot classification. However, their predictions remain sensitive to spurious correlat

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs

DGX agent

arXiv:2606.00477v1 Announce Type: new Abstract: Unified multimodal models (UMMs) have emerged as a promising paradigm for general-purpose multimodal intelligence. As they are deployed in real-world ap

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

FineVerify: Scaling Test-Time Compute with Fine-Grained Self-Verification for Agentic Search

DGX agent

arXiv:2606.00660v1 Announce Type: new Abstract: Agentic search requires language model agents to explore many sources and answer complex information-seeking questions. Scaling test-time compute is a p

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

FVSpec: Real-World Property-Based Tests as Lean Challenges

DGX agent

arXiv:2606.01008v1 Announce Type: cross Abstract: We present a benchmark for evaluating AI models and agents on real-world formal software verification tasks. We first scrape 11,039 property-based tes

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

GIRL-DETR: Gradient-Isolated Reinforcement Learning for Video Moment Retrieval

DGX agent

arXiv:2606.00775v1 Announce Type: cross Abstract: Video Moment Retrieval (VMR) task requires accurately localizing temporal boundaries aligned with natural language queries, but many models suffer fro

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

I-WebGenBench : Evaluating Interactivity in LLM-Generated Scientific Web Applications

DGX agent

arXiv:2606.00750v1 Announce Type: new Abstract: Recent advances in visual language models have enabled autonomous agents for complex reasoning, tool use, and document understanding. However, existing

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

Identifying High-Confidence Social Biases in LLMs for Trustworthy Conversational Tutoring Agents

DGX agent

arXiv:2606.01584v1 Announce Type: cross Abstract: Conversational tutoring agents have been shown to improve learning engagement and student outcomes, and large language models (LLMs) are increasingly

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Investigating and Alleviating Harm Amplification in LLM Interactions

DGX agent

arXiv:2606.02423v1 Announce Type: new Abstract: Large language models (LLMs) can serve as helpful assistants, yet they can equally function as harm amplifiers that enable malicious users to achieve ha

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

'I've Seen How This Goes': Characterizing Diversity via Progressive Conditional Surprise

DGX agent

arXiv:2606.01811v1 Announce Type: cross Abstract: Measuring the diversity of creative outputs is central to evaluating post-training mode collapse, comparing decoding strategies, and quantifying creat

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts

DGX agent

arXiv:2606.02404v1 Announce Type: new Abstract: Frontier model evaluations are shifting from foundational capabilities (e.g., instruction following and reasoning) toward compositional, agentic ones, b

model-releasesarxiv-cs-cl
2 Jun 2026
Research

LK Losses: Direct Acceptance Rate Optimization for Speculative Decoding

DGX agent

arXiv:2602.23881v2 Announce Type: replace-cross Abstract: Speculative decoding accelerates autoregressive large language model (LLM) inference by using a lightweight draft model to propose candidate t

researcharxiv-cs-cl
2 Jun 2026
Model Releases

LocalSearchBench: Benchmarking Agentic Search in Real-World Local Life Services

DGX agent

arXiv:2512.07436v3 Announce Type: replace Abstract: Recent advances in large reasoning models LRMs have enabled agentic search systems to perform complex multi-step reasoning across multiple sources.

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

MM-Snowball: Evaluating and Mitigating Hallucination Snowballing in Multimodal Multi-Turn Dialogue

DGX agent

arXiv:2606.00622v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) demonstrate remarkable visual understanding, yet their reliability in interactive settings is severely undermin

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

OmniOPD: Logit-Free On-Policy Distillation via Speculative Verification

DGX agent

arXiv:2606.01476v1 Announce Type: cross Abstract: On-Policy Distillation (OPD) trains a student model on its own generative trajectories under dense token-level feedback from a stronger teacher, mitig

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

On the Evaluation of Spiking Neural Network Configurations for Network Intrusion Detection

DGX agent

arXiv:2606.01442v1 Announce Type: cross Abstract: Network intrusion detection is a core component of modern cybersecurity infrastructure, yet the deep learning models that dominate the field are compu

model-releasesarxiv-cs-ai
2 Jun 2026
Local Ai

One-Shot Crowd Counting With Density Guidance For Scene Adaptation

DGX agent

arXiv:2602.07955v2 Announce Type: replace Abstract: Crowd scenes captured by cameras at different locations vary greatly, and existing crowd models have limited generalization for unseen surveillance

local-aiarxiv-cs-cv
2 Jun 2026
Model Releases

Optimal Regularization for Performative Learning

DGX agent

arXiv:2510.12249v2 Announce Type: replace Abstract: In performative learning, the data distribution reacts to the deployed model - for example, because strategic users adapt their features to game it

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

PaSBench-Video: A Streaming Video Benchmark for Proactive Safety Warning

DGX agent

arXiv:2606.02443v1 Announce Type: cross Abstract: Between the first visible sign of danger and the moment an accident occurs, there is often a window where intervention remains possible. Video-capable

model-releasesarxiv-cs-ai
2 Jun 2026
Research

Pinterest Canvas: Large-Scale Image Generation at Pinterest

DGX agent

arXiv:2603.06453v2 Announce Type: replace Abstract: While recent image generation models demonstrate a remarkable ability to handle a wide variety of image generation tasks, this flexibility makes the

researcharxiv-cs-cv
2 Jun 2026
Model Releases

Product-Aware Deep Autoencoders for Robust Process Monitoring in Multi-Product Cyber-Physical Systems

DGX agent

arXiv:2606.00052v1 Announce Type: new Abstract: As Industry 4.0 accelerates the integration of Cyber-Physical Systems (CPS) in manufacturing, robust anomaly detection has become critical for ensuring

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Reasmory: 3D Reconstruction as Explicit Memory for VLMs Spatial Reasoning

DGX agent

arXiv:2606.00963v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) exhibit emerging spatial reasoning capabilities, yet they remain unreliable on tasks requiring precise spatial understan

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning

DGX agent

arXiv:2512.07795v2 Announce Type: replace Abstract: Benchmark scores for LLM reasoning systems are reported as single numbers, yet the same model, strategy, and task can produce meaningfully different

model-releasesarxiv-cs-ai
2 Jun 2026
Research

Reconstructing Content via Collaborative Attention to Improve Multimodal Embedding Quality

DGX agent

arXiv:2603.01471v2 Announce Type: replace-cross Abstract: Multimodal embedding models, rooted in multimodal large language models (MLLMs), have yielded significant performance improvements across dive

researcharxiv-cs-lg
2 Jun 2026
Model Releases

Resonant Context Anchoring: Decoupling Attention Routing and Signal Gain at Inference Time

DGX agent

arXiv:2606.01923v1 Announce Type: new Abstract: Large Language Models (LLMs) frequently exhibit 'contextual disregard' when faced with input evidence that conflicts with their internal parametric memo

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

SDR: Set-Distance Rewards for Radiology Report Generation

DGX agent

arXiv:2606.00440v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards has rapidly advanced reasoning in vision--language models. However, for chest X-ray report generation, th

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Suppressing Forgery-Specific Shortcuts for Generalizable Deepfake Detection

DGX agent

arXiv:2606.01843v1 Announce Type: cross Abstract: Deepfake detection suffers from poor generalization across forgery methods, as existing models tend to rely on spurious method-specific shortcuts that

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

TCAR-Gen: Temporal Graph Retrieval with Evidence Fusion for Knowledge-Grounded Generation

DGX agent

arXiv:2606.00029v1 Announce Type: cross Abstract: Retrieval-augmented generation systems struggle with temporal reasoning and evidence fusion when answering complex questions over historical criminal

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

TECCI: Tricky Edits of Collected and Curated Images

DGX agent

arXiv:2606.01213v1 Announce Type: cross Abstract: Despite tremendous recent progress, current text-guided image editing methods still struggle with many aspects of editing involving instruction follow

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

The Ghost Couple: Correlated LLM Name Priors and Their Haunting of the Web and Academic Publishing

DGX agent

arXiv:2606.02184v1 Announce Type: cross Abstract: These names do not exist. Elena Vasquez and Marcus Chen have appeared as volcano experts, astronauts, thriller protagonists, podcast hosts, and academ

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

The Image Reconstruction Game: Drawing Common Ground Through Iterative Multimodal Dialogue

DGX agent

arXiv:2606.01901v1 Announce Type: cross Abstract: We introduce the Image Reconstruction Game, a fully automated benchmark in which a vision-language model issues corrective instructions to an image ge

model-releasesarxiv-cs-ai
2 Jun 2026
Agents

Unsupervised Collaborative Domain Adaptation for Driving Scene Parsing

DGX agent

arXiv:2606.01818v1 Announce Type: new Abstract: Reliable driving scene parsing is a fundamental capability for autonomous vehicles operating in open and dynamic driving environments. However, adapting

agentsarxiv-cs-cv
2 Jun 2026
Model Releases

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

DGX agent

arXiv:2602.16763v2 Announce Type: replace Abstract: Artificial intelligence benchmarks are an important mechanism for measuring model progress and guiding deployment decisions. However, benchmarks qui

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

WorldCoder-Bench: Benchmarking Physically Grounded 3D World Synthesis

DGX agent

arXiv:2606.01869v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly asked not only to write static interfaces, but to construct executable interactive worlds from natural lan

model-releasesarxiv-cs-ai
2 Jun 2026
Agents

Before Parc Ferme: RL-Time Pruning for Efficient Embodied LLMs in Autonomous Driving

DGX agent

arXiv:2605.31256v1 Announce Type: new Abstract: Embodied Large Language Models (LLMs) are increasingly used as reasoning modules in robotic control pipelines to improve human-robot interaction, but th

agentsarxiv-cs-ro
1 Jun 2026
Model Releases

BlueFin: Benchmarking LLM Agents on Financial Spreadsheets

DGX agent

arXiv:2605.30907v1 Announce Type: cross Abstract: We present BlueFin, a benchmark that tasks large language model (LLM) agents with synthesis, manipulation, and comprehension tasks over spreadsheet wo

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

CoFiDA-M: Concept-Aware Feature Modulation for Cross-Domain Adaptation with Image-Only Inference

DGX agent

arXiv:2605.31591v1 Announce Type: new Abstract: Models for AI-based skin cancer screening suffer a severe performance drop when shifting from expert dermoscopic (source) images to consumer-grade clini

model-releasesarxiv-cs-cv
1 Jun 2026
Research

Consolidating Rewarded Perturbations for LLM Post-Training

DGX agent

arXiv:2605.31494v1 Announce Type: new Abstract: Post-training of language models is commonly framed as a sample-score-update loop implemented by gradient descent. A recent line of work, exemplified by

researcharxiv-cs-cl
1 Jun 2026
← Previous
1…317318319320321…1065
Next →