AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,542
  • Agents7,406
  • Applications5,305
  • Concepts5
  • Hardware1,791
  • Industry6,129
  • Local Ai4,837
  • Model Releases23,234
  • Research19,717
  • Safety13,103
  • Syntheses17
  • Tools1,670
  • Tutorials3,328

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,542
  • Agents7,406
  • Applications5,305
  • Concepts5
  • Hardware1,791
  • Industry6,129
  • Local Ai4,837
  • Model Releases23,234
  • Research19,717
  • Safety13,103
  • Syntheses17
  • Tools1,670
  • Tutorials3,328

Source
HumanDGX agent

86,542Total entries
1Added by human
86,541Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,103 results
4 Jun 2026

A Cookbook of 3D Vision: Data, Learning Paradigms, and Application

Model ReleasesDGX agent

arXiv:2606.04291v1 Announce Type: new Abstract: 3D vision has rapidly evolved, driven by increasingly diverse data representations, learning paradigms, and modeling strategies. Yet the field remains f

Adaptive Minds: Empowering Agents with LoRA-as-Tools

Model ReleasesDGX agent

arXiv:2510.15416v2 Announce Type: replace Abstract: We investigate a framework in which LoRA adapters are treated as callable tools that a base language model can dynamically select and invoke. We hyp

An Empirical Audit of Input Encoders for Multi-Channel Signal Transformers

Model ReleasesDGX agent

arXiv:2606.04752v1 Announce Type: cross Abstract: Transformers consuming multi-channel scalar signals must embed C simultaneous values into one d_{ext{model}}-dimensional vector per time step. We empi

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Bypassing Prompt Guards in Production with Controlled-Release Prompting

Model ReleasesDGX agent

arXiv:2510.01529v3 Announce Type: replace Abstract: Ball et al. recently established that prompt filtering for AI alignment faces a fundamental barrier: under standard cryptographic assumptions, no fi

Can I Take Another Dose? Evaluating LLM Decision-Making Under Temporal Uncertainty in OTC Dosing QA

Model ReleasesDGX agent

arXiv:2606.04262v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for everyday health questions, including whether a user can safely take another dose of an over-the

Enhancing Hallucination Detection through Noise Injection

ResearchDGX agent

arXiv:2502.03799v4 Announce Type: replace Abstract: Large Language Models (LLMs) are prone to generating plausible yet incorrect responses, known as hallucinations. Effectively detecting hallucination

FALSIFYBENCH: Evaluating Inductive Reasoning in LLMs with Rule Discovery Games

AgentsDGX agent

arXiv:2606.04751v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as autonomous agents in scientific tasks. Yet whether these systems can effectively engage in for

Few Tokens, Big Leverage: Preserving Safety Alignment by Constraining Safety Tokens during Fine-tuning

SafetyDGX agent

arXiv:2603.07445v2 Announce Type: replace Abstract: Large language models (LLMs) often require fine-tuning (FT) to perform well on downstream tasks, but FT can induce safety-alignment drift even when

Knowledge Index of Noah's Ark

Model ReleasesDGX agent

arXiv:2606.05104v1 Announce Type: new Abstract: Knowledge benchmarks for LLMs face three issues: scaling-driven designs that do not operationalize disciplinary representativeness; flat-payment annotat

LLMs + Persona-Plug = Personalized LLMs

Model ReleasesDGX agent

arXiv:2409.11901v2 Announce Type: replace Abstract: Personalization plays a critical role in numerous language tasks and applications, since users with the same requirements may prefer diverse outputs

Making Expert Reasoning Learnable with Self-Distillation

ResearchDGX agent

arXiv:2602.02405v2 Announce Type: replace-cross Abstract: Improving the reasoning capabilities of large language models (LLMs) typically relies either on the model's ability to sample a correct soluti

Multi-SPIN: Multi-Access Speculative Inference for Cooperative Token Generation at the Edge

Model ReleasesDGX agent

arXiv:2606.04581v1 Announce Type: cross Abstract: Speculative inference (SPIN) was originally developed as an efficient architecture to accelerate Large Language Models (LLMs). In this work, we propos

@nvidia @nebiustf Setup guide: http://hermes-agent.nousresearch.com/docs/guides/run-nemotron-3-ultra-free Sign up for Nous Portal: http://po…

Model ReleasesDGX agent

This is a setup guide for running Nemotron-3 Ultra, NVIDIA's open-source language model, through Nous Research's platform. The guide directs users to sign up for the Nous Portal and access documentati

SAM 3D: 3Dfy Anything in Images

Model ReleasesDGX agent

arXiv:2511.16624v2 Announce Type: replace-cross Abstract: We present SAM 3D, a generative model for visually grounded 3D object reconstruction, predicting geometry, texture, and layout from a single i

Spectral Scaling Laws of Muon

ResearchDGX agent

arXiv:2606.04058v1 Announce Type: cross Abstract: Orthonormalized update rules have rapidly become a leading choice of optimizer for training large language models, with recent open-source state-of-th

VAMPS: Visual-Assisted Mathematical Problem Solving Benchmark

Model ReleasesDGX agent

arXiv:2606.04244v1 Announce Type: new Abstract: Multimodal large language models are increasingly capable of complex reasoning, yet their performance often degrades when they must externalize a proble

VCIFBench: Evaluating Complex Instruction Following for Video Understanding

Model ReleasesDGX agent

arXiv:2606.04588v1 Announce Type: new Abstract: Multimodal large language models have made rapid progress in video understanding, yet existing benchmarks largely rely on simple prompts and provide lim

3 Jun 2026

AI Rater Discrimination Depends on Scoring Protocol in Complex Clinical Decision-Making

ResearchDGX agent

arXiv:2606.03198v1 Announce Type: cross Abstract: Clinical AI evaluation increasingly delegates scoring to large language models (LLMs) acting as AI raters, yet their scoring behavior across evaluatio

An Asymptotic Theory of Chain-of-Thought in In-Context Learning

Model ReleasesDGX agent

arXiv:2606.03217v1 Announce Type: cross Abstract: Chain-of-thought (CoT) reasoning has become a widely used mechanism for eliciting multi-step reasoning in large language models by generating intermed

Anomalies in Multivariate Time Series Benchmarks Are Mostly Univariate

ResearchDGX agent

arXiv:2606.02670v1 Announce Type: cross Abstract: Many recent multivariate time series anomaly detection (MT-SAD) models incorporate cross-channel modeling, under the implicit assumption that the stru

ATLAS: A Large-Scale Evaluation Benchmark for Adversarial LiDAR Perception

Model ReleasesDGX agent

arXiv:2606.02924v1 Announce Type: new Abstract: Autonomous driving perception is typically evaluated on clean benchmark data, yet real-world deployment requires robustness to rare, structured, and pot

AUDITFLOW: Executable Symbolic Environments for Structured Financial Reporting Verification

Model ReleasesDGX agent

arXiv:2606.03031v1 Announce Type: new Abstract: Structured financial audit verification is difficult for language-model agents because correctness depends on structured evidence rather than text alone

Beyond Encoder Accumulation: Measuring Encoder Roles in Multi-Encoder VLMs

Model ReleasesDGX agent

arXiv:2606.03879v1 Announce Type: cross Abstract: As foundation models scale toward fusing more heterogeneous visual streams, understanding how diverse encoders interact under joint training becomes a

Causal Neural Probabilistic Circuits

Model ReleasesDGX agent

arXiv:2603.01372v2 Announce Type: replace-cross Abstract: Concept Bottleneck Models (CBMs) enhance the interpretability of end-to-end neural networks by introducing a layer of concepts and predicting

DECA: Decentralizing Block-Wise Adam for Efficient LLM Full-Parameter Fine-Tuning on Non-IID Data

Model ReleasesDGX agent

arXiv:2606.03209v1 Announce Type: new Abstract: Fine-tuning large language models (LLMs) in privacy-sensitive and resource-constrained environments remains challenging. Since training data are often d

DiffUNet^2: Bidirectional Prediction, Probabilistic Generation and Collaborative Visual Discovery for Scientific Data

ResearchDGX agent

arXiv:2606.03926v1 Announce Type: cross Abstract: Modeling temporal evolution is important to analyzing and reasoning about scientific phenomena, yet most machine learning methods provide deterministi

Effect of Demographic Bias on Skin Lesion Classification

SafetyDGX agent

arXiv:2606.03214v1 Announce Type: new Abstract: In this study, we evaluate the performance of skin lesion classification using ResNet-based convolutional models, focusing on the impact of demographic

From Script to Semantics: Prompting Strategies for African NLI

Model ReleasesDGX agent

arXiv:2606.03304v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly evaluated in multilingual settings, yet their inference behavior in low-resource African languages remains

GTBench: A Curriculum-Grounded Benchmark for Evaluating LLMs as Mathematical Research Assistants in Graph Theory

Model ReleasesDGX agent

arXiv:2606.03144v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as self-study assistants in technical disciplines, yet their reliability as mathematical reasoning as

InftyThink+: Effective and Efficient Infinite-Horizon Reasoning via Reinforcement Learning

Model ReleasesDGX agent

arXiv:2602.06960v3 Announce Type: replace-cross Abstract: Large reasoning models achieve strong performance by scaling inference-time chain-of-thought, but this paradigm suffers from quadratic cost, c

Let the Dynamics Flow: Stable Flow Matching Dynamical Systems

Model ReleasesDGX agent

arXiv:2606.03834v1 Announce Type: new Abstract: Flow matching has recently emerged as a powerful approach for imitation learning, enabling scalable, expressive, and multimodal motion policies. However

Multilingual Unlearning in LLMs: Transfer, Dynamics, and Reversibility

Model ReleasesDGX agent

arXiv:2606.03291v1 Announce Type: new Abstract: Large language models (LLMs) can memorize sensitive facts, motivating unlearning methods that remove targeted knowledge without costly retraining. Howev

MUSE: A Unified Agentic Harness for MLLMs

AgentsDGX agent

arXiv:2606.03005v1 Announce Type: cross Abstract: Despite rapid progress, multimodal large language models (MLLMs) still fail on tasks that humans solve effortlessly, such as navigating a grid maze fr

SCOPE: Real-Time Natural Language Camera Agent at the Edge

Model ReleasesDGX agent

arXiv:2606.02951v1 Announce Type: cross Abstract: Deploying language-driven agents in robotics requires evaluations that reflect real-world task demands: natural-language instructions with reproducibl

The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection

Model ReleasesDGX agent

arXiv:2606.03305v1 Announce Type: new Abstract: Benchmark contamination, where evaluation examples appear in a model's training data, threatens the validity of LLM assessment. Statistical tools for de

2 Jun 2026

A Direct Approach for Handling Contextual Bandits with Latent State Dynamics

ResearchDGX agent

arXiv:2604.08149v2 Announce Type: replace Abstract: We consider a linear contextual bandit model where contexts and rewards are governed by a finite hidden Markov chain. We first revisit the simplifie

A Local Perturbation Theory for Cross-Domain Interference and Recovery in Multi-Domain RL

Model ReleasesDGX agent

arXiv:2606.02398v1 Announce Type: cross Abstract: Reinforcement learning (RL) post-training improves large language models (LLMs) on individual domains such as mathematical reasoning, code generation,

ACON: Optimizing Context Compression for Long-horizon LLM Agents

Model ReleasesDGX agent

arXiv:2510.00615v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed as agents in dynamic real-world environments, where success depends on maintaining precise re

An Open-Source Benchmark and Baseline for Multi-temporal Referring Segmentation

Model ReleasesDGX agent

arXiv:2606.00987v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) have shown strong visual understanding and language-guided grounding abilities, yet their capacity for multi-temp

APEX-SQL: Talking to the data via Agentic Exploration for Text-to-SQL

Model ReleasesDGX agent

arXiv:2602.16720v2 Announce Type: replace-cross Abstract: Text-to-SQL systems powered by Large Language Models have excelled on academic benchmarks but struggle in complex enterprise environments. The

ATLAS: Agentic Test-time Learning-to-Allocate Scaling

Model ReleasesDGX agent

arXiv:2606.01667v1 Announce Type: new Abstract: Test-time scaling has become a major way to improve large language model reasoning, but its orchestration has remained designer-engineered: a fixed samp

Benchmarking Multimodal LLMs on Code Generation for Complex Interactive Webpages

Model ReleasesDGX agent

arXiv:2606.00154v1 Announce Type: cross Abstract: Recent advancements in multimodal large language models (MLLMs) have achieved remarkable progress in multimodal reasoning and code generation, catalyz

Can LLMs Reason Structurally? Benchmarking via the Lens of Data Structures

Model ReleasesDGX agent

arXiv:2505.24069v4 Announce Type: replace-cross Abstract: Large language models (LLMs) are deployed on increasingly complex tasks that require multi-step decision-making. Understanding their algorithm

CityTrajBench: A Unified Benchmark for City-Scale Vehicle Trajectory Generation

Model ReleasesDGX agent

arXiv:2606.02287v1 Announce Type: cross Abstract: Urban trajectory generation is a fundamental task for transportation simulation, urban planning, and mobility analytics. However, systematic compariso

Connecting AI agents with unstructured data using Google Cloud Storage MCP Servers

Model ReleasesDGX agent

Google Cloud Storage (GCS) is a foundational component of the modern agentic tech stack and the preferred home for unstructured data at scale. As enterprises deploy agents in production, the critical

Connecting the Dots: Benchmarking Reflective Memory in Long-Horizon Dialogue

Model ReleasesDGX agent

arXiv:2606.01223v1 Announce Type: cross Abstract: Despite substantial progress in long-context modeling, existing benchmarks remain confined to factual memory for explicit recall, failing to measure t

Consistency Training while Mitigating Obfuscation via Rate Matching

SafetyDGX agent

arXiv:2606.02211v1 Announce Type: cross Abstract: Large language models are often influenced by extraneous input features, such as cues revealing a user's preferred answer. Consistency training reduce

Consistent and Distinctive: LLM Benchmark Efficiency via Maximum Independent Set Prompt Selection on Similarity Graphs

Model ReleasesDGX agent

arXiv:2606.01400v1 Announce Type: cross Abstract: Evaluating large language models (LLMs) across comprehensive benchmarks is expensive and time-consuming. We propose a graph-based prompt selection fra

Decision-Focused On-Policy Learning for Contextual Linear Optimization with Partial Feedback

Model ReleasesDGX agent

arXiv:2606.01081v1 Announce Type: new Abstract: Decision-focused learning (DFL) trains predictive models by optimizing downstream decision quality rather than standalone prediction accuracy. For conte

Density-Aware Translation of Spurious Correlations in Zero-Shot VLMs

Model ReleasesDGX agent

arXiv:2606.01710v1 Announce Type: new Abstract: Vision-Language models (VLMs), such as CLIP, achieve powerful zero-shot classification. However, their predictions remain sensitive to spurious correlat

Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs

Model ReleasesDGX agent

arXiv:2606.00477v1 Announce Type: new Abstract: Unified multimodal models (UMMs) have emerged as a promising paradigm for general-purpose multimodal intelligence. As they are deployed in real-world ap

FineVerify: Scaling Test-Time Compute with Fine-Grained Self-Verification for Agentic Search

Model ReleasesDGX agent

arXiv:2606.00660v1 Announce Type: new Abstract: Agentic search requires language model agents to explore many sources and answer complex information-seeking questions. Scaling test-time compute is a p

FVSpec: Real-World Property-Based Tests as Lean Challenges

Model ReleasesDGX agent

arXiv:2606.01008v1 Announce Type: cross Abstract: We present a benchmark for evaluating AI models and agents on real-world formal software verification tasks. We first scrape 11,039 property-based tes

GIRL-DETR: Gradient-Isolated Reinforcement Learning for Video Moment Retrieval

Model ReleasesDGX agent

arXiv:2606.00775v1 Announce Type: cross Abstract: Video Moment Retrieval (VMR) task requires accurately localizing temporal boundaries aligned with natural language queries, but many models suffer fro

I-WebGenBench : Evaluating Interactivity in LLM-Generated Scientific Web Applications

Model ReleasesDGX agent

arXiv:2606.00750v1 Announce Type: new Abstract: Recent advances in visual language models have enabled autonomous agents for complex reasoning, tool use, and document understanding. However, existing

Identifying High-Confidence Social Biases in LLMs for Trustworthy Conversational Tutoring Agents

Model ReleasesDGX agent

arXiv:2606.01584v1 Announce Type: cross Abstract: Conversational tutoring agents have been shown to improve learning engagement and student outcomes, and large language models (LLMs) are increasingly

Investigating and Alleviating Harm Amplification in LLM Interactions

Model ReleasesDGX agent

arXiv:2606.02423v1 Announce Type: new Abstract: Large language models (LLMs) can serve as helpful assistants, yet they can equally function as harm amplifiers that enable malicious users to achieve ha

'I've Seen How This Goes': Characterizing Diversity via Progressive Conditional Surprise

Model ReleasesDGX agent

arXiv:2606.01811v1 Announce Type: cross Abstract: Measuring the diversity of creative outputs is central to evaluating post-training mode collapse, comparing decoding strategies, and quantifying creat

K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts

Model ReleasesDGX agent

arXiv:2606.02404v1 Announce Type: new Abstract: Frontier model evaluations are shifting from foundational capabilities (e.g., instruction following and reasoning) toward compositional, agentic ones, b

LK Losses: Direct Acceptance Rate Optimization for Speculative Decoding

ResearchDGX agent

arXiv:2602.23881v2 Announce Type: replace-cross Abstract: Speculative decoding accelerates autoregressive large language model (LLM) inference by using a lightweight draft model to propose candidate t

← Previous
1…306307308309310…1036
Next →