AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,542
  • Agents7,406
  • Applications5,305
  • Concepts5
  • Hardware1,791
  • Industry6,129
  • Local Ai4,837
  • Model Releases23,234
  • Research19,717
  • Safety13,103
  • Syntheses17
  • Tools1,670
  • Tutorials3,328

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,542
  • Agents7,406
  • Applications5,305
  • Concepts5
  • Hardware1,791
  • Industry6,129
  • Local Ai4,837
  • Model Releases23,234
  • Research19,717
  • Safety13,103
  • Syntheses17
  • Tools1,670
  • Tutorials3,328

Source
HumanDGX agent

86,542Total entries
1Added by human
86,541Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,103 results
2 Jun 2026

LocalSearchBench: Benchmarking Agentic Search in Real-World Local Life Services

Model ReleasesDGX agent

arXiv:2512.07436v3 Announce Type: replace Abstract: Recent advances in large reasoning models LRMs have enabled agentic search systems to perform complex multi-step reasoning across multiple sources.

MM-Snowball: Evaluating and Mitigating Hallucination Snowballing in Multimodal Multi-Turn Dialogue

Model ReleasesDGX agent

arXiv:2606.00622v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) demonstrate remarkable visual understanding, yet their reliability in interactive settings is severely undermin

OmniOPD: Logit-Free On-Policy Distillation via Speculative Verification

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.01476v1 Announce Type: cross Abstract: On-Policy Distillation (OPD) trains a student model on its own generative trajectories under dense token-level feedback from a stronger teacher, mitig

On the Evaluation of Spiking Neural Network Configurations for Network Intrusion Detection

Model ReleasesDGX agent

arXiv:2606.01442v1 Announce Type: cross Abstract: Network intrusion detection is a core component of modern cybersecurity infrastructure, yet the deep learning models that dominate the field are compu

One-Shot Crowd Counting With Density Guidance For Scene Adaptation

Local AiDGX agent

arXiv:2602.07955v2 Announce Type: replace Abstract: Crowd scenes captured by cameras at different locations vary greatly, and existing crowd models have limited generalization for unseen surveillance

Optimal Regularization for Performative Learning

Model ReleasesDGX agent

arXiv:2510.12249v2 Announce Type: replace Abstract: In performative learning, the data distribution reacts to the deployed model - for example, because strategic users adapt their features to game it

PaSBench-Video: A Streaming Video Benchmark for Proactive Safety Warning

Model ReleasesDGX agent

arXiv:2606.02443v1 Announce Type: cross Abstract: Between the first visible sign of danger and the moment an accident occurs, there is often a window where intervention remains possible. Video-capable

Pinterest Canvas: Large-Scale Image Generation at Pinterest

ResearchDGX agent

arXiv:2603.06453v2 Announce Type: replace Abstract: While recent image generation models demonstrate a remarkable ability to handle a wide variety of image generation tasks, this flexibility makes the

Product-Aware Deep Autoencoders for Robust Process Monitoring in Multi-Product Cyber-Physical Systems

Model ReleasesDGX agent

arXiv:2606.00052v1 Announce Type: new Abstract: As Industry 4.0 accelerates the integration of Cyber-Physical Systems (CPS) in manufacturing, robust anomaly detection has become critical for ensuring

Reasmory: 3D Reconstruction as Explicit Memory for VLMs Spatial Reasoning

Model ReleasesDGX agent

arXiv:2606.00963v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) exhibit emerging spatial reasoning capabilities, yet they remain unreliable on tasks requiring precise spatial understan

ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning

Model ReleasesDGX agent

arXiv:2512.07795v2 Announce Type: replace Abstract: Benchmark scores for LLM reasoning systems are reported as single numbers, yet the same model, strategy, and task can produce meaningfully different

Reconstructing Content via Collaborative Attention to Improve Multimodal Embedding Quality

ResearchDGX agent

arXiv:2603.01471v2 Announce Type: replace-cross Abstract: Multimodal embedding models, rooted in multimodal large language models (MLLMs), have yielded significant performance improvements across dive

Resonant Context Anchoring: Decoupling Attention Routing and Signal Gain at Inference Time

Model ReleasesDGX agent

arXiv:2606.01923v1 Announce Type: new Abstract: Large Language Models (LLMs) frequently exhibit 'contextual disregard' when faced with input evidence that conflicts with their internal parametric memo

SDR: Set-Distance Rewards for Radiology Report Generation

Model ReleasesDGX agent

arXiv:2606.00440v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards has rapidly advanced reasoning in vision--language models. However, for chest X-ray report generation, th

Suppressing Forgery-Specific Shortcuts for Generalizable Deepfake Detection

Model ReleasesDGX agent

arXiv:2606.01843v1 Announce Type: cross Abstract: Deepfake detection suffers from poor generalization across forgery methods, as existing models tend to rely on spurious method-specific shortcuts that

TCAR-Gen: Temporal Graph Retrieval with Evidence Fusion for Knowledge-Grounded Generation

Model ReleasesDGX agent

arXiv:2606.00029v1 Announce Type: cross Abstract: Retrieval-augmented generation systems struggle with temporal reasoning and evidence fusion when answering complex questions over historical criminal

TECCI: Tricky Edits of Collected and Curated Images

Model ReleasesDGX agent

arXiv:2606.01213v1 Announce Type: cross Abstract: Despite tremendous recent progress, current text-guided image editing methods still struggle with many aspects of editing involving instruction follow

The Ghost Couple: Correlated LLM Name Priors and Their Haunting of the Web and Academic Publishing

Model ReleasesDGX agent

arXiv:2606.02184v1 Announce Type: cross Abstract: These names do not exist. Elena Vasquez and Marcus Chen have appeared as volcano experts, astronauts, thriller protagonists, podcast hosts, and academ

The Image Reconstruction Game: Drawing Common Ground Through Iterative Multimodal Dialogue

Model ReleasesDGX agent

arXiv:2606.01901v1 Announce Type: cross Abstract: We introduce the Image Reconstruction Game, a fully automated benchmark in which a vision-language model issues corrective instructions to an image ge

Unsupervised Collaborative Domain Adaptation for Driving Scene Parsing

AgentsDGX agent

arXiv:2606.01818v1 Announce Type: new Abstract: Reliable driving scene parsing is a fundamental capability for autonomous vehicles operating in open and dynamic driving environments. However, adapting

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

Model ReleasesDGX agent

arXiv:2602.16763v2 Announce Type: replace Abstract: Artificial intelligence benchmarks are an important mechanism for measuring model progress and guiding deployment decisions. However, benchmarks qui

WorldCoder-Bench: Benchmarking Physically Grounded 3D World Synthesis

Model ReleasesDGX agent

arXiv:2606.01869v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly asked not only to write static interfaces, but to construct executable interactive worlds from natural lan

1 Jun 2026

Before Parc Ferme: RL-Time Pruning for Efficient Embodied LLMs in Autonomous Driving

AgentsDGX agent

arXiv:2605.31256v1 Announce Type: new Abstract: Embodied Large Language Models (LLMs) are increasingly used as reasoning modules in robotic control pipelines to improve human-robot interaction, but th

BlueFin: Benchmarking LLM Agents on Financial Spreadsheets

Model ReleasesDGX agent

arXiv:2605.30907v1 Announce Type: cross Abstract: We present BlueFin, a benchmark that tasks large language model (LLM) agents with synthesis, manipulation, and comprehension tasks over spreadsheet wo

CoFiDA-M: Concept-Aware Feature Modulation for Cross-Domain Adaptation with Image-Only Inference

Model ReleasesDGX agent

arXiv:2605.31591v1 Announce Type: new Abstract: Models for AI-based skin cancer screening suffer a severe performance drop when shifting from expert dermoscopic (source) images to consumer-grade clini

Consolidating Rewarded Perturbations for LLM Post-Training

ResearchDGX agent

arXiv:2605.31494v1 Announce Type: new Abstract: Post-training of language models is commonly framed as a sample-score-update loop implemented by gradient descent. A recent line of work, exemplified by

CSULoRA: Closest Safe Update Low-Rank Adaptation

Model ReleasesDGX agent

arXiv:2605.30640v1 Announce Type: cross Abstract: Low-rank adaptation has become a standard method for parameter-efficient fine-tuning of large language models, but even small amounts of unsafe or adv

Effective Reasoning Chains Reduce Intrinsic Dimensionality

Model ReleasesDGX agent

arXiv:2602.09276v2 Announce Type: replace-cross Abstract: Chain-of-thought (CoT) reasoning and its variants have substantially improved the performance of language models on complex reasoning tasks, y

Eigenvectors of Experts are Training-free Non-collapsing Routers

TutorialsDGX agent

arXiv:2605.30992v1 Announce Type: new Abstract: Sparse Mixture of Experts (SMoE) architectures improve the training efficiency of Large Language Models (LLMs) by routing input tokens to a selected sub

EUDAIMONIA: Evaluating Undesirable Dynamics in AI

Model ReleasesDGX agent

arXiv:2605.30654v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as conversational partners for companionship, emotional disclosure, and interpersonal advice, but t

Evaluating using Mock Tool Calls to Quarantine Untrusted Prompt Inputs

ResearchDGX agent

arXiv:2605.30521v1 Announce Type: new Abstract: Large language models must frequently process untrusted inputs, such as judging an answer from another model or running tasks like spam and harm classif

FBHM: Functional Benchmarking and Steering of VLMs for Hateful Meme Detection

Model ReleasesDGX agent

arXiv:2605.31349v1 Announce Type: cross Abstract: Hateful meme detection remains a formidable challenge for vision-language models, as existing benchmarks are structurally observational - confounding

HERMES: Towards Efficient and Verifiable Mathematical Reasoning in LLMs

Model ReleasesDGX agent

arXiv:2511.18760v2 Announce Type: replace Abstract: Informal mathematics has been central to modern large language model (LLM) reasoning, offering flexibility and efficient construction of arguments.

LegSegNet: A Public Deep Learning System for Lower Extremity CT Tissue Segmentation and Quantification

Model ReleasesDGX agent

arXiv:2605.30829v1 Announce Type: new Abstract: Lower extremity computed tomography (CT) contains clinically relevant information for body composition analysis, sarcopenia assessment, and musculoskele

MineExplorer: Evaluating Open-World Exploration of MLLM Agents in Minecraft

Model ReleasesDGX agent

arXiv:2605.30931v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have shown strong capabilities in perception, reasoning, and action generation. However, their ability to susta

Nemotron 3 Ultra: Frontier smart. 5X faster. 30% cheaper. 💚💚💚

Model ReleasesDGX agent

Nemotron 3 Ultra is NVIDIA's latest language model featuring significant improvements in speed (5X faster) and cost efficiency (30% cheaper) compared to previous versions, positioning it as a frontier

PhyDrawGen: Physically Grounded Diagram Generation from Natural Language

Model ReleasesDGX agent

arXiv:2605.30512v1 Announce Type: new Abstract: Generating physics diagrams from text requires strict adherence to physical laws. While current generative models produce visually plausible outputs, th

Post-Training LLMs as Better Decision-Making Agents: A Regret-Minimization Approach

ResearchDGX agent

arXiv:2511.04393v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed as 'agents' for decision-making (DM) in interactive and dynamic environments. Yet, since they

PRISM: Progressive Reasoning through Iterative Slot Memory for Vision

Model ReleasesDGX agent

arXiv:2605.30942v1 Announce Type: new Abstract: Modern vision models process images in a single feed-forward pass, which limits their ability to recover missing evidence or refine uncertain representa

Re-examining Low Rank adaptation for private LLM fine-tuning

Model ReleasesDGX agent

arXiv:2510.01137v3 Announce Type: replace Abstract: Privacy is a central concern when fine-tuning large language models (LLMs) on sensitive data, and differentially private stochastic gradient descent

Safe Equilibrium Policy Optimization for Strategic Agent Policies

Model ReleasesDGX agent

arXiv:2605.30854v1 Announce Type: cross Abstract: Language models fine-tuned with reinforcement learning typically optimize for task reward, ignoring multi-agent strategic structure. Because these age

SERA: Soft-Verified Efficient Repository Agents

Model ReleasesDGX agent

arXiv:2601.20789v3 Announce Type: replace Abstract: Open-weight coding agents should hold a fundamental advantage over closed-source systems because they can specialize to private codebases, encoding

Smaller and Faster 3DGS via Post-Training Dictionary Learning

Model ReleasesDGX agent

arXiv:2605.30396v1 Announce Type: cross Abstract: 3D Gaussian Splatting (3DGS) is a promising neural scene representation for real-time rendering, but trained models often suffer from large memory foo

Speculative Decoding Across Languages

ResearchDGX agent

arXiv:2605.30580v1 Announce Type: new Abstract: Speculative decoding has become a crucial component of large language model (LLM) inference, enabling faster generation by drafting multiple tokens and

Symbolic Intermediaries as a Linguistic-Numerical Interface for LLM-Driven Geometric Reasoning

Model ReleasesDGX agent

arXiv:2505.17607v3 Announce Type: replace Abstract: Large Language Models (LLMs) display reasoning capabilities over linguistic and symbolic objects but have limited capabilities to directly interpret

Triaging Threats to Specialized Guardrails

Model ReleasesDGX agent

arXiv:2605.30693v1 Announce Type: cross Abstract: Building robust safety guardrails is essential for deploying Large Language Models across diverse real-world applications. However, this goal remains

Who Gets Credit or Blame? Attributing Accountability in Modern AI Systems

SafetyDGX agent

arXiv:2506.00175v5 Announce Type: replace-cross Abstract: Modern AI systems are typically developed through multiple stages-pretraining, fine-tuning rounds, and subsequent adaptation or alignment, whe

29 May 2026

Apertus LLM Family Expansion via Distillation and Quantization

Model ReleasesDGX agent

arXiv:2605.29128v1 Announce Type: new Abstract: The wide adoption of LLMs has led to their use in great variety of applications and scenarios, such as chatbot assistants and data annotation, creating

CLUBench: A Clustering Benchmark

Model ReleasesDGX agent

arXiv:2605.29933v1 Announce Type: new Abstract: Clustering is a fundamental problem in data science with a long-standing research history, yielding numerous insightful algorithms. Despite this progres

ConMoE: Expert-Pool Consolidation via Prototype Reassignment for MoE Compression

Model ReleasesDGX agent

arXiv:2605.29350v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) language models reduce per-token computation but still require storing and serving all experts, making deployment memory-intens

DiffSpot: Can VLMs Spot Fine-Grained Visual Differences in Web Interfaces?

Model ReleasesDGX agent

arXiv:2605.29615v1 Announce Type: cross Abstract: Vision-language models (VLMs) have made strong progress on high-level image-text alignment, yet their ability to perceive subtle visual differences re

EVL-ECG: Efficient ECG Interpretation With Multi-Aspect Heterogeneous Knowledge Distillation

Model ReleasesDGX agent

arXiv:2605.29977v1 Announce Type: new Abstract: High-fidelity ECG interpretation is increasingly reliant on massive foundation models, yet their deployment in clinical edge-care remains hindered by ex

Evolutionary Dynamics of Cooperation in Next-Generation LLM Agent Systems: A Cross-Provider Empirical Extension

Model ReleasesDGX agent

arXiv:2605.29874v1 Announce Type: cross Abstract: Do next-generation LLM agents inherit the cooperative biases documented in their predecessors, or does scale and provider diversity reshape equilibriu

Fairness Beyond Demographics: Optimizing Performance Across Appearance-Based Hidden Cohorts in Medical Imaging

Model ReleasesDGX agent

arXiv:2605.29827v1 Announce Type: new Abstract: Medical image analysis models can exhibit performance disparities across patient subgroups, threatening clinical safety and fairness. Existing methods t

Gram: Assessing sabotage propensities via automated alignment auditing

Model ReleasesDGX agent

arXiv:2605.30322v1 Announce Type: cross Abstract: We introduce Gram, an automated alignment auditing framework to assess the propensity of AI agents to engage in sabotage. We evaluate Gemini models ac

Horizon Activation Mapping for Neural Networks in Time Series Forecasting

ResearchDGX agent

arXiv:2601.02094v4 Announce Type: replace Abstract: Neural networks for time series forecasting have relied on error metrics and architecture-specific interpretability approaches for model selection t

Knowledge Offloading: Decomposing LLMs into Sparse Backbones and Memory Modules

Model ReleasesDGX agent

arXiv:2605.29075v1 Announce Type: new Abstract: LLMs encode both general capabilities and domain-specific knowledge in a single set of parameters. We ask whether this capacity can be reorganized: keep

LFQ: Logit-aware Final-block Quantization for Boosting the Generation Quality of Low-Bit Quantized LLMs

ResearchDGX agent

arXiv:2605.29756v1 Announce Type: new Abstract: As large language models continue to scale, low-bit weight-only post-training quantization (PTQ) offers a practical solution to their memory-efficient d

LLUMI: Improving LLM Writing Assistance for Mental Health Support with Online Community Feedback

SafetyDGX agent

arXiv:2605.30273v1 Announce Type: cross Abstract: Large language models (LLMs) show promise in generating supportive responses for mental health queries, but improving their usefulness, empathy, and s

LUMINA: A Multi-Vendor Mammography Benchmark with Energy Harmonization Protocol

Model ReleasesDGX agent

arXiv:2603.14644v3 Announce Type: replace-cross Abstract: Publicly available full-field digital mammography (FFDM) datasets remain limited in size, clinical annotations, and vendor diversity, hinderin

← Previous
1…307308309310311…1036
Next →