AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent
84,648Total entries
1Added by human
84,647Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,620 results
1 Jun 2026

Convergence of Two-Timescale Markovian Stochastic Approximations with Applications in Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.31172v1 Announce Type: new Abstract: This work studies the convergence of two-timescale stochastic approximations (SA), a class of iterative algorithms that update two sets of parameters in

Count Anything

Model ReleasesDGX agent

arXiv:2605.30846v1 Announce Type: new Abstract: Object counting remains fragmented across domain-specific datasets and task formulations, despite rapid progress in generalist vision models. Existing c

Counterfactual Trace Auditing of LLM Agent Skills

Model ReleasesDGX agent

arXiv:2605.11946v2 Announce Type: replace Abstract: Large Language Model agents are increasingly augmented with agent skills. Current evaluation methods for skills remain limited. Most deployed benchm


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs

Model ReleasesDGX agent

arXiv:2605.30611v1 Announce Type: cross Abstract: Scientific figures are among the most effective means of communicating complex research ideas, yet producing publication-quality illustrations remains

CSULoRA: Closest Safe Update Low-Rank Adaptation

Model ReleasesDGX agent

arXiv:2605.30640v1 Announce Type: cross Abstract: Low-rank adaptation has become a standard method for parameter-efficient fine-tuning of large language models, but even small amounts of unsafe or adv

DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories

Model ReleasesDGX agent

arXiv:2602.10809v2 Announce Type: replace Abstract: Existing multimodal retrieval systems excel at semantic matching but implicitly assume that query-image relevance can be measured in isolation. This

DeMaVLA: A Vision-Language-Action Foundation Model for Generalizable Deformable Manipulation

Model ReleasesDGX agent

arXiv:2605.31286v1 Announce Type: cross Abstract: Real-world household robots require Vision-Language-Action (VLA) foundation models that can acquire reusable manipulation skills across diverse object

Depth-Dependent Indirect Prompt Injection in Tool-Calling ReAct Agents: Injection Depth, Payload Framing, and Turn-Budget Sensitivity

Model ReleasesDGX agent

arXiv:2605.30686v1 Announce Type: cross Abstract: ReAct agents that interleave chain-of-thought reasoning with tool calls are increasingly deployed for real tasks such as scheduling, file retrieval, a

Design and Evaluation of Multi-Agent AI Oracle Systems for Prediction Market Resolution

Model ReleasesDGX agent

arXiv:2605.30802v1 Announce Type: cross Abstract: Prediction markets aggregate collective intelligence to forecast uncertain events, but their utility depends on reliable outcome resolution. Existing

Dex2HOI: Dexterous Bimanual Two-Object Interaction Generation

Model ReleasesDGX agent

arXiv:2605.30444v1 Announce Type: new Abstract: Recent advances in 4D Human-Object Interaction (HOI) generation have enabled increasingly realistic motion synthesis, particularly for single-object man

Distilling Counterfactual Reasoning from Language to Vision: Causal Graph Guided Post-Training for Video Understanding

Model ReleasesDGX agent

arXiv:2511.19923v2 Announce Type: replace-cross Abstract: Vision Language Models (VLMs) have recently shown significant advancements in video understanding, especially in feature alignment, event reas

Diving into Kronecker Adapters: Component Design Matters

Model ReleasesDGX agent

arXiv:2602.01267v2 Announce Type: replace Abstract: Kronecker adapters have emerged as a promising approach for fine-tuning large-scale models, enabling high-rank updates through tunable component str

Do covariates explain why these groups differ? The choice of reference group can reverse conclusions in the Oaxaca-Blinder decomposition

Model ReleasesDGX agent

arXiv:2603.29972v2 Announce Type: replace-cross Abstract: Scientists often want to explain why an outcome is different in two groups. For instance, differences in patient mortality rates across two ho

DriveMA: Driving Vision-Language-Action Models with verifiable Meta-Actions

Model ReleasesDGX agent

arXiv:2605.31271v1 Announce Type: new Abstract: Driving Vision-Language-Action Models (Driving VLAs) aim to use language to improve end-to-end planning, but the language-action gap limits this promise

DTBench: A Synthetic Benchmark for Document-to-Table Extraction

Model ReleasesDGX agent

arXiv:2602.13812v3 Announce Type: replace-cross Abstract: Document-to-table (Doc2Table) extraction derives structured tables from unstructured documents under a target schema, enabling reliable and ve

DTG-Restore: Training-Free Diffusion Refinement for Generative Video Super-Resolution

Model ReleasesDGX agent

arXiv:2605.30431v1 Announce Type: new Abstract: Recent progress in video diffusion models has enabled remarkable generative fidelity, yet leveraging these priors for restoration remains limited by the

DynaTree: Dynamic Agentic Retrieval Tree for Time-Sensitive News Retrieval

Model ReleasesDGX agent

arXiv:2605.31377v1 Announce Type: cross Abstract: Agentic Retrieval-Augmented Generation improves retrieval by integrating planning, tool use, and iterative reasoning, but existing agentic RAG methods

Effective Reasoning Chains Reduce Intrinsic Dimensionality

Model ReleasesDGX agent

arXiv:2602.09276v2 Announce Type: replace-cross Abstract: Chain-of-thought (CoT) reasoning and its variants have substantially improved the performance of language models on complex reasoning tasks, y

EGOSTREAM: A Diagnostic Benchmark for Streaming Episodic Memory in Egocentric Vision

Model ReleasesDGX agent

arXiv:2605.31557v1 Announce Type: new Abstract: Continuous episodic memory is a core capability for autonomous agents operating in dynamic, real-world environments, yet current streaming video benchma

EHRBench: An Automated and Reliable EHR-based Benchmark for Clinical Decision Making with LLMs

Model ReleasesDGX agent

arXiv:2605.30637v1 Announce Type: new Abstract: Clinical decision-making (CDM) is central to real-world clinical workflows, where clinicians infer diagnoses, select treatments, or anticipate future he

ElasticMem: Latent Memory as a Learnable Resource for LLM Agents

Model ReleasesDGX agent

arXiv:2605.30690v1 Announce Type: new Abstract: Long-term memory is essential for LLM agents to reason coherently across extended interactions, personalize responses, and reuse past experience. Howeve

EMBGuard: Constructing Hazard-Aware Guardrails for Safe Planning in Embodied Agents

Model ReleasesDGX agent

arXiv:2605.30924v1 Announce Type: new Abstract: MLLM-powered embodied agents deployed in real-world environments encounter physical hazards. However, existing approaches lack explicit mechanisms for i

Emergent Languages in Populations of Language Model Agents: From Token Efficiency to Oversight Evasion

Model ReleasesDGX agent

arXiv:2605.31170v1 Announce Type: cross Abstract: Monitoring autonomous language model agents currently relies mostly on surface behavior. But what happens when agent populations invent new languages

End-to-End Compression for Tabular Foundation Models

Model ReleasesDGX agent

arXiv:2602.05649v2 Announce Type: replace Abstract: The long-standing dominance of gradient-boosted decision trees for tabular data has recently been challenged by in-context learning tabular foundati

ERGeoBench:A Comprehensive Benchmark for Embodied Reasoning and Geo-localization in Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2605.31251v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have shown strong potential as embodied agents, yet embodied geo-localization remains underexplored due to th

Error Amplification Limits ANN-to-SNN Conversion in Continuous Control

Model ReleasesDGX agent

arXiv:2601.21778v2 Announce Type: replace-cross Abstract: Spiking Neural Networks (SNNs) can achieve competitive performance by converting already existing well-trained Artificial Neural Networks (ANN

EUDAIMONIA: Evaluating Undesirable Dynamics in AI

Model ReleasesDGX agent

arXiv:2605.30654v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as conversational partners for companionship, emotional disclosure, and interpersonal advice, but t

Evaluating Factual Density in Multi-Source RAG: A Study in Medical AI Accuracy

Model ReleasesDGX agent

arXiv:2605.31506v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) is the current industry standard for grounding AI in real-world facts. Traditional retrieval methods rely on keyw

EvoDefense: Co-Evolving Black-Box Defense with Large Language Models

Model ReleasesDGX agent

arXiv:2605.31140v1 Announce Type: cross Abstract: Large Language Models (LLMs) remain highly vulnerable to diverse attacks, particularly in black-box settings where the internals of target models are

Expand Neurons, Not Parameters

Model ReleasesDGX agent

arXiv:2510.04500v2 Announce Type: replace Abstract: This work demonstrates how increasing the number of neurons in a network without increasing its total number of non-zero parameters improves perform

Expert Merging in Sparse Mixture of Experts with Nash Bargaining

Model ReleasesDGX agent

arXiv:2510.16138v2 Announce Type: replace Abstract: Existing expert merging strategies for Sparse Mixture of Experts (SMoE) typically rely on input-dependent or input-independent averaging of expert p

ExpGraph: Model-Agnostic Experience Learning with Graph-Structured Memory for LLM Agents

Model ReleasesDGX agent

arXiv:2605.30712v1 Announce Type: new Abstract: Large language model (LLM) agents have shown strong capabilities in reasoning, tool use, and multi-step interaction, but they often solve tasks from scr

Exploring Autonomous Agentic Data Engineering for Model Specialization

Model ReleasesDGX agent

arXiv:2605.30407v1 Announce Type: cross Abstract: Large Language Models (LLMs) have demonstrated strong performance on general tasks, while often struggling to adapt to specialized domains without hig

Extending AI for Research to the Humanities: A Multi-Agent Framework for Evidence-Grounded Scholarship

Model ReleasesDGX agent

arXiv:2605.30947v1 Announce Type: new Abstract: LLM-based research agents have advanced rapidly in science and engineering, where research is organized around executable experiments, code, and quantit

Eywa: Provenance-Grounded Long-Term Memory for AI Agents

Model ReleasesDGX agent

arXiv:2605.30771v1 Announce Type: new Abstract: AI agents that persist across sessions need memory they can retrieve, audit, update, and erase. Existing memory systems often collapse source evidence,

FAM-Bench: A Multimodal Benchmark for Condition-Aware Food-as-Medicine Reasoning

Model ReleasesDGX agent

arXiv:2605.31410v1 Announce Type: new Abstract: Food-as-Medicine requires models to reason beyond what a dish is or what nutrition it contains: they must decide whether a concrete food choice is appro

FBHM: Functional Benchmarking and Steering of VLMs for Hateful Meme Detection

Model ReleasesDGX agent

arXiv:2605.31349v1 Announce Type: cross Abstract: Hateful meme detection remains a formidable challenge for vision-language models, as existing benchmarks are structurally observational - confounding

FEM-Bench: A Structured Scientific Reasoning Benchmark for Evaluating Code-Generating LLMs

Model ReleasesDGX agent

arXiv:2512.20732v2 Announce Type: replace-cross Abstract: As LLMs advance their reasoning capabilities about the physical world, the absence of rigorous benchmarks for evaluating their ability to gene

Fingerprint launches AI Assistant Detection to spot traffic from ChatGPT, Gemini and Claude

Model ReleasesDGX agent

Device intelligence company FingerprintJS Inc. today launched a preview of two products built to identify traffic from artificial intelligence assistants, addressing a detection gap that has opened as

Float8@2bits: Entropy Coding Enables Data-Free Model Compression

Model ReleasesDGX agent

arXiv:2601.22787v2 Announce Type: replace Abstract: Post-training compression is currently divided into two contrasting regimes. On the one hand, fast, data-free, and model-agnostic methods (e.g., NF4

FOCUS: Forcing In-Context Object Localization through Visual Support Constraints and Policy Optimization

Model ReleasesDGX agent

arXiv:2605.31145v1 Announce Type: cross Abstract: In-context localization (ICL) seeks to localize a target object specified by a small set of support examples in a query image, operating on the fly wi

ForecastCompass: Guiding Agentic Forecasting with Adaptive Factor Memory

Model ReleasesDGX agent

arXiv:2605.30858v1 Announce Type: new Abstract: Agentic forecasting is important for decision-making in dynamic environments, but it remains challenging because agents must reason from incomplete, tim

From Prompt Injection to Persistent Control: Defending Agentic Harness Against Trojan Backdoors

Model ReleasesDGX agent

arXiv:2605.31042v1 Announce Type: cross Abstract: LLM agents are evolving from conversational chatbots to operational tools in real-world workspaces. In local agentic harnesses, an LLM can read and wr

From Weak Cues to Real Identities: Evaluating Inference-Driven De-Anonymization in LLM Agents

Model ReleasesDGX agent

arXiv:2603.18382v2 Announce Type: replace Abstract: Anonymization is often assumed to protect privacy once explicit identifiers are removed, because re-identification has historically required special

FSM-Net: An Efficient Frequency-Spatial Network for Real-World Deblurring

Model ReleasesDGX agent

arXiv:2605.31400v1 Announce Type: new Abstract: Real-world image deblurring demands both high-fidelity restoration and computational efficiency, a balance existing methods often struggle to achieve. I

Gait2Hip-60: A Unified Deep Learning Benchmark for Predicting Hip Muscle Forces and Joint Moments from Multi-Cadence Gait Kinematics

Model ReleasesDGX agent

arXiv:2605.30374v1 Announce Type: new Abstract: Estimating hip muscle forces and joint moments during gait typically relies on musculoskeletal simulation, which is informative but time-consuming and d

GEM-Bench: A Benchmark for Ad-Injected Response Generation within Generative Engine Marketing

Model ReleasesDGX agent

arXiv:2509.14221v3 Announce Type: replace-cross Abstract: Generative Engine Marketing (GEM) is an emerging ecosystem for monetizing generative engines, such as LLM-based chatbots, by seamlessly integr

Gemini’s new AI agent is about as good as Google’s demo

Model ReleasesDGX agent

Google's new '24/7' AI agent, Gemini Spark, can be shockingly good at doing things on your behalf. But I'm not sure it's worth the financial cost and potential privacy tradeoffs. The company gave me a

Generalistic or Specific Embeddings, Which is Better? An Empirical Study on Search for Clinical Coding in Non-English Languages

Model ReleasesDGX agent

arXiv:2605.30529v1 Announce Type: cross Abstract: Sentence-embedding models for semantic search are overwhelmingly developed and evaluated on English corpora. When applied to clinical retrieval in oth

Generating Graph-like Rules for Knowledge Graph Reasoning via Diffusion Models

Model ReleasesDGX agent

arXiv:2605.30747v1 Announce Type: new Abstract: Logical rules constitute a cornerstone of knowledge graph (KG) reasoning, valued for their interpretability and ability to model relational patterns. Ho

Generating Reports or Repeating Templates? Measuring and Mitigating Template Collapse in 3D CT Report Generation

Model ReleasesDGX agent

arXiv:2605.30984v1 Announce Type: cross Abstract: Modern 3D medical vision-language models (VLMs) can generate fluent radiology-style text while exhibit critically low pathology detection and output d

Go-UT-Bench: A Fine-Tuning Dataset for LLM-Based Unit Test Generation in Go

Model ReleasesDGX agent

arXiv:2511.10868v2 Announce Type: replace Abstract: Training data imbalance poses a major challenge for code LLMs. Most available data heavily over represents raw opensource code while underrepresenti

Goldfish: Monolingual Language Models for 350 Languages

Model ReleasesDGX agent

arXiv:2408.10441v3 Announce Type: replace Abstract: For many low-resource languages, the only available language models are large multilingual models trained on many languages simultaneously. Despite

Graph-Conditioned Mixture of Graph Neural Network Experts for Traffic Forecasting

Model ReleasesDGX agent

arXiv:2605.30486v1 Announce Type: cross Abstract: Spatio-temporal forecasting on sensor graphs is commonly tackled with a single backbone architecture applied uniformly across all nodes, although grap

GraphARC: A Comprehensive Benchmark for Graph-Based Abstract Reasoning

Model ReleasesDGX agent

arXiv:2605.31031v1 Announce Type: new Abstract: Relational reasoning lies at the heart of intelligence, but existing benchmarks are typically confined to formats such as grids or text. We introduce Gr

GUI-C^2: Coarse-to-Fine GUI Grounding via Difficulty-Aware Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.30884v1 Announce Type: new Abstract: Existing agentic reinforcement learning methods for GUI grounding have limitations at two levels. At the data level, current approaches typically treat

Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents

Model ReleasesDGX agent

arXiv:2605.30621v1 Announce Type: new Abstract: LLM agents are increasingly deployed as systems built around editable external harnesses, including prompts, skills, memories and tools, that shape task

Hedging on the Frontier: Learning New Tasks with Few Samples

Model ReleasesDGX agent

arXiv:2605.30997v1 Announce Type: cross Abstract: When a learner faces a new task with few samples, it must leverage any available side information. In practice, this often comes in the form of model

HERMES: Towards Efficient and Verifiable Mathematical Reasoning in LLMs

Model ReleasesDGX agent

arXiv:2511.18760v2 Announce Type: replace Abstract: Informal mathematics has been central to modern large language model (LLM) reasoning, offering flexibility and efficient construction of arguments.

How Far Can You Grow? Characterizing the Extrapolation Frontier of Graph Generative Models for Materials Science

Model ReleasesDGX agent

arXiv:2602.09309v2 Announce Type: replace-cross Abstract: Every generative model for crystalline materials harbors a critical structure size beyond which its outputs become unreliable; we call this th

← Previous
1…188189190191192…377
Next →