AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,648Total entries
1Added by human
84,647Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
17,288 results
Model Releases

Auditing LLM Benchmarks with Item Response Theory

DGX agent

arXiv:2605.30504v1 Announce Type: new Abstract: LLM benchmark labels are frozen at release and silently propagated into downstream benchmarks, errors and all. We introduce an Item Response Theory-base

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Auto-Discovery-Bench: Diagnosing Structured State Tracking in Oracle-Guided Discovery

DGX agent

arXiv:2502.15224v2 Announce Type: replace-cross Abstract: Interactive discovery requires agents to maintain and update structured beliefs over many rounds of feedback. Before evaluating agents in nois

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Automated Prediction of Postoperative Pancreatic Fistula Using Preoperative Computed Tomography

DGX agent

arXiv:2605.31539v1 Announce Type: new Abstract: Postoperative pancreatic fistula (POPF) is a serious complication after pancreatic resection, increasing morbidity, hospital stay, and healthcare costs.

model-releasesarxiv-cs-cv
1 Jun 2026
Model Releases

Automating Formal Verification with Reinforcement Learning and Recursive Inference

DGX agent

arXiv:2605.30914v1 Announce Type: new Abstract: Automated formal verification remains challenging for large language models because data for proof assistants and verification-aware languages is scarce

model-releasesarxiv-cs-lg
1 Jun 2026
Model Releases

Balanced LoRA: Removing Parameter Invariance to Accelerate Convergence

DGX agent

arXiv:2605.31484v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) is the most widely adopted method for fine-tuning large language models. Notably, LoRA is inherently overparameterized: multi

model-releasesarxiv-cs-lg
1 Jun 2026
Model Releases

Bandwidth Allocation with Device Partitioning for Federated Learning over Industrial IoT networks

DGX agent

arXiv:2605.30892v1 Announce Type: new Abstract: We consider a federated learning (FL) system in which Industrial Internet-of-Things (IIoT) devices collaboratively train a global model over wireless ch

model-releasesarxiv-cs-lg
1 Jun 2026
Model Releases

Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education

DGX agent

arXiv:2605.31212v1 Announce Type: cross Abstract: AI systems are increasingly used to support educational content creation, yet it remains unclear whether they can generate outputs that faithfully rep

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Benchmarking Uncertainty and its Disentanglement in multi-label Chest X-Ray Classification

DGX agent

arXiv:2508.04457v2 Announce Type: replace-cross Abstract: Reliable uncertainty quantification is crucial for trustworthy decision-making and the deployment of AI models in medical imaging. While prior

model-releasesarxiv-cs-lg
1 Jun 2026
Model Releases

BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali

DGX agent

arXiv:2605.31483v1 Announce Type: new Abstract: Despite Bengali being the sixth most spoken language in the world, no prior work has systematically evaluated hallucination in large language models (LL

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

Beyond Agreement: Scoring Panel-Surfaced Biomedical Entity Candidates for Curator Triage

DGX agent

arXiv:2605.30826v1 Announce Type: cross Abstract: Biomedical NER is deceptively simple for modern LLMs: plausible biomedical mentions are easy to surface, but corpus-convention correctness depends on

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Beyond ReLU: Bifurcation, Oversmoothing, and Topological Priors

DGX agent

arXiv:2602.15634v2 Announce Type: replace Abstract: Graph Neural Networks (GNNs) learn node representations through iterative network-based message-passing. While powerful, deep GNNs suffer from overs

model-releasesarxiv-cs-lg
1 Jun 2026
Model Releases

Beyond Static Dialogues: Benchmarking Realistic, Heterogeneous, and Evolving Long-Term Memory

DGX agent

arXiv:2605.31086v1 Announce Type: new Abstract: In existing memory benchmarks for Large Language Models (LLMs), the evaluated dialogue sessions often lack long-term semantic consistency, and the under

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

BilliardPhys-Bench: Benchmarking Physical Reasoning and Visual Dynamics of Multimodal LLMs

DGX agent

arXiv:2605.30900v1 Announce Type: new Abstract: Current multimodal models handle static image recognition well, but intuitive physical reasoning remains a weakness. Predicting how objects will move an

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

BlueFin: Benchmarking LLM Agents on Financial Spreadsheets

DGX agent

arXiv:2605.30907v1 Announce Type: cross Abstract: We present BlueFin, a benchmark that tasks large language model (LLM) agents with synthesis, manipulation, and comprehension tasks over spreadsheet wo

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

BOKBO (Best of K Bad Options): Calibrated Abstention for VLA Policies

DGX agent

arXiv:2605.30660v1 Announce Type: new Abstract: Test-time scaling for vision-language-action (VLA) policies, methods such as RoboMonkey, SEAL, MG-Select, and V-GPS, samples K candidate action chunks a

model-releasesarxiv-cs-lg
1 Jun 2026
Model Releases

Bottom-up Policy Optimization: Your Language Model Policy Secretly Contains Internal Policies

DGX agent

arXiv:2512.19673v3 Announce Type: replace-cross Abstract: Existing reinforcement learning (RL) approaches treat large language models (LLMs) as a unified policy, overlooking their internal mechanisms.

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Bounded Behavioral Indistinguishability for Black-Box LLM Distillation

DGX agent

arXiv:2605.30448v1 Announce Type: cross Abstract: Black-box LLM distillation is usually evaluated as an output-matching problem: a student is considered successful when its responses are semantically

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

Breaking the Simplification Bottleneck in Amortized Neural Symbolic Regression

DGX agent

arXiv:2602.08885v5 Announce Type: replace-cross Abstract: Symbolic regression (SR) aims to discover interpretable analytical expressions that accurately describe observed data. Amortized SR promises t

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Calibrated Preference Learning: The Case of Label Ranking

DGX agent

arXiv:2605.30447v1 Announce Type: cross Abstract: Calibration, the alignment of predicted probabilities with true outcome frequencies, is essential for reliable decision-making. While extensively stud

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Can LLM Teams Play What? Where? When?

DGX agent

arXiv:2605.30459v1 Announce Type: new Abstract: Large language models (LLMs) remain limited on tasks requiring indirect reasoning, cultural knowledge, and coordinated hypothesis testing. We investigat

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

Can Subgraph Explanations Be Weaponized to Steal Graph Neural Networks?

DGX agent

arXiv:2605.30470v1 Announce Type: new Abstract: Graph Machine Learning as a Service (GMLaaS) platforms increasingly implement explainability interfaces to meet regulatory transparency requirements. Ho

model-releasesarxiv-cs-lg
1 Jun 2026
Model Releases

CanLegalRAGBench: Evaluating Retrieval-Augmented Generation on Canadian Case Law

DGX agent

arXiv:2605.30497v1 Announce Type: new Abstract: RAG-based legal assistants have been growing in popularity, but LLM hallucinations remain a key issue and potentially undermines justice. While benchmar

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

Caspar: CUDA Accelerator for Symbolic Programming with Adaptive Reordering

DGX agent

arXiv:2605.30583v1 Announce Type: new Abstract: We present Caspar, a library that makes the power of modern GPUs more accessible in robotics and provides a state-of-the-art nonlinear GPU solver that c

model-releasesarxiv-cs-ro
1 Jun 2026
Model Releases

Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

DGX agent

arXiv:2503.08679v5 Announce Type: replace Abstract: Recent studies indicate that when faced with explicit biases in prompts, models often omit mentioning these biases in their Chain-of-Thought (CoT) o

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

CodeGolf Bench: A Multi-Language Benchmark for Evaluating Concise Code Generation Capabilities of Large Language Models

DGX agent

arXiv:2605.30394v1 Announce Type: cross Abstract: This paper introduces Code Bench, a benchmark capable of evaluating Large Language Models (LLMs) concise code generation abilities in 60 programming l

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

CoFiDA-M: Concept-Aware Feature Modulation for Cross-Domain Adaptation with Image-Only Inference

DGX agent

arXiv:2605.31591v1 Announce Type: new Abstract: Models for AI-based skin cancer screening suffer a severe performance drop when shifting from expert dermoscopic (source) images to consumer-grade clini

model-releasesarxiv-cs-cv
1 Jun 2026
Model Releases

Conditional Coverage Diagnostics for Conformal Prediction

DGX agent

arXiv:2512.11779v2 Announce Type: replace-cross Abstract: Evaluating conditional coverage remains one of the most persistent challenges in assessing the reliability of predictive systems. Although con

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

ConTrans: Learning Text-enhanced Local-global Temporal Representations for Zero-shot Temporal Action Localization

DGX agent

arXiv:2605.30689v1 Announce Type: cross Abstract: Zero-shot Temporal Action Localization (ZS-TAL) aims to detect and locate previously unseen actions in untrimmed videos. However, existing approaches

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Convergence of Steepest Descent and Adam under Non-Uniform Smoothness

DGX agent

arXiv:2605.30648v1 Announce Type: new Abstract: Recent work has analyzed the convergence of first-order methods under non-uniform smoothness assumptions that better model the loss landscape in machine

model-releasesarxiv-cs-lg
1 Jun 2026
Model Releases

Convergence of Two-Timescale Markovian Stochastic Approximations with Applications in Reinforcement Learning

DGX agent

arXiv:2605.31172v1 Announce Type: new Abstract: This work studies the convergence of two-timescale stochastic approximations (SA), a class of iterative algorithms that update two sets of parameters in

model-releasesarxiv-cs-lg
1 Jun 2026
Model Releases

Count Anything

DGX agent

arXiv:2605.30846v1 Announce Type: new Abstract: Object counting remains fragmented across domain-specific datasets and task formulations, despite rapid progress in generalist vision models. Existing c

model-releasesarxiv-cs-cv
1 Jun 2026
Model Releases

Counterfactual Trace Auditing of LLM Agent Skills

DGX agent

arXiv:2605.11946v2 Announce Type: replace Abstract: Large Language Model agents are increasingly augmented with agent skills. Current evaluation methods for skills remain limited. Most deployed benchm

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs

DGX agent

arXiv:2605.30611v1 Announce Type: cross Abstract: Scientific figures are among the most effective means of communicating complex research ideas, yet producing publication-quality illustrations remains

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

CSULoRA: Closest Safe Update Low-Rank Adaptation

DGX agent

arXiv:2605.30640v1 Announce Type: cross Abstract: Low-rank adaptation has become a standard method for parameter-efficient fine-tuning of large language models, but even small amounts of unsafe or adv

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories

DGX agent

arXiv:2602.10809v2 Announce Type: replace Abstract: Existing multimodal retrieval systems excel at semantic matching but implicitly assume that query-image relevance can be measured in isolation. This

model-releasesarxiv-cs-cv
1 Jun 2026
Model Releases

DeMaVLA: A Vision-Language-Action Foundation Model for Generalizable Deformable Manipulation

DGX agent

arXiv:2605.31286v1 Announce Type: cross Abstract: Real-world household robots require Vision-Language-Action (VLA) foundation models that can acquire reusable manipulation skills across diverse object

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Depth-Dependent Indirect Prompt Injection in Tool-Calling ReAct Agents: Injection Depth, Payload Framing, and Turn-Budget Sensitivity

DGX agent

arXiv:2605.30686v1 Announce Type: cross Abstract: ReAct agents that interleave chain-of-thought reasoning with tool calls are increasingly deployed for real tasks such as scheduling, file retrieval, a

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Design and Evaluation of Multi-Agent AI Oracle Systems for Prediction Market Resolution

DGX agent

arXiv:2605.30802v1 Announce Type: cross Abstract: Prediction markets aggregate collective intelligence to forecast uncertain events, but their utility depends on reliable outcome resolution. Existing

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Dex2HOI: Dexterous Bimanual Two-Object Interaction Generation

DGX agent

arXiv:2605.30444v1 Announce Type: new Abstract: Recent advances in 4D Human-Object Interaction (HOI) generation have enabled increasingly realistic motion synthesis, particularly for single-object man

model-releasesarxiv-cs-cv
1 Jun 2026
Model Releases

Distilling Counterfactual Reasoning from Language to Vision: Causal Graph Guided Post-Training for Video Understanding

DGX agent

arXiv:2511.19923v2 Announce Type: replace-cross Abstract: Vision Language Models (VLMs) have recently shown significant advancements in video understanding, especially in feature alignment, event reas

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

Diving into Kronecker Adapters: Component Design Matters

DGX agent

arXiv:2602.01267v2 Announce Type: replace Abstract: Kronecker adapters have emerged as a promising approach for fine-tuning large-scale models, enabling high-rank updates through tunable component str

model-releasesarxiv-cs-lg
1 Jun 2026
Model Releases

Do covariates explain why these groups differ? The choice of reference group can reverse conclusions in the Oaxaca-Blinder decomposition

DGX agent

arXiv:2603.29972v2 Announce Type: replace-cross Abstract: Scientists often want to explain why an outcome is different in two groups. For instance, differences in patient mortality rates across two ho

model-releasesarxiv-cs-lg
1 Jun 2026
Model Releases

DriveMA: Driving Vision-Language-Action Models with verifiable Meta-Actions

DGX agent

arXiv:2605.31271v1 Announce Type: new Abstract: Driving Vision-Language-Action Models (Driving VLAs) aim to use language to improve end-to-end planning, but the language-action gap limits this promise

model-releasesarxiv-cs-cv
1 Jun 2026
Model Releases

DTBench: A Synthetic Benchmark for Document-to-Table Extraction

DGX agent

arXiv:2602.13812v3 Announce Type: replace-cross Abstract: Document-to-table (Doc2Table) extraction derives structured tables from unstructured documents under a target schema, enabling reliable and ve

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

DTG-Restore: Training-Free Diffusion Refinement for Generative Video Super-Resolution

DGX agent

arXiv:2605.30431v1 Announce Type: new Abstract: Recent progress in video diffusion models has enabled remarkable generative fidelity, yet leveraging these priors for restoration remains limited by the

model-releasesarxiv-cs-cv
1 Jun 2026
Model Releases

DynaTree: Dynamic Agentic Retrieval Tree for Time-Sensitive News Retrieval

DGX agent

arXiv:2605.31377v1 Announce Type: cross Abstract: Agentic Retrieval-Augmented Generation improves retrieval by integrating planning, tool use, and iterative reasoning, but existing agentic RAG methods

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Effective Reasoning Chains Reduce Intrinsic Dimensionality

DGX agent

arXiv:2602.09276v2 Announce Type: replace-cross Abstract: Chain-of-thought (CoT) reasoning and its variants have substantially improved the performance of language models on complex reasoning tasks, y

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

EGOSTREAM: A Diagnostic Benchmark for Streaming Episodic Memory in Egocentric Vision

DGX agent

arXiv:2605.31557v1 Announce Type: new Abstract: Continuous episodic memory is a core capability for autonomous agents operating in dynamic, real-world environments, yet current streaming video benchma

model-releasesarxiv-cs-cv
1 Jun 2026
← Previous
1…174175176177178…361
Next →