AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,563 results
11 May 2026

Beyond the Black Box: Interpretability of Agentic AI Tool Use

Model ReleasesDGX agent

arXiv:2605.06890v1 Announce Type: new Abstract: AI agents are promising for high-stakes enterprise workflows, but dependable deployment remains limited because tool-use failures are difficult to diagn

BGM-IV: an AI-powered Bayesian generative modeling approach for instrumental variable analysis

Model ReleasesDGX agent

arXiv:2605.07029v1 Announce Type: cross Abstract: Instrumental-variable (IV) regression enables causal estimation under endogeneity, but modern IV problems often involve nonlinear structural effects a

Bi3: A Biplatform, Bicultural, Biperson Dataset for Social Robot Navigation

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.06863v1 Announce Type: new Abstract: We contribute Bi3, a dataset of social robot navigation among groups of people in a constrained lab space. Compared to prior data collection efforts for

BioProVLA-Agent: An Affordable, Protocol-Driven, Vision-Enhanced VLA-Enabled Embodied Multi-Agent System with Closed-Loop-Capable Reasoning for Biological Laboratory Manipulation

Model ReleasesDGX agent

arXiv:2605.07306v1 Announce Type: cross Abstract: Biological laboratory automation can reduce repetitive manual work and improve reproducibility, but reliable embodied execution in wet-lab environment

bispectrum: Selective G-Bispectra Made Practical

Model ReleasesDGX agent

arXiv:2605.07270v1 Announce Type: new Abstract: Many machine learning tasks are invariant under the action of a group G of transformations: signal classification can be invariant under translations, i

BoHA: Blockwise Hadamard Product Adaptation for Parameter-Efficient Fine-Tuning

Model ReleasesDGX agent

arXiv:2509.21637v2 Announce Type: replace Abstract: Parameter-efficient fine-tuning (PEFT) of large language models trains a small task-specific parameter set while keeping the pretrained model frozen

Breaking Spatial Uniformity: Prior-Guided Mamba with Radial Serialization for Lens Flare Removal

Model ReleasesDGX agent

arXiv:2605.07650v1 Announce Type: new Abstract: Lens flares, caused by complex optical aberrations, severely degrade image quality especially in nighttime photography. Although recent restoration meth

BRIDGE: Background Routing and Isolated Discrete Gating for Coarse-Mask Local Editing

Model ReleasesDGX agent

arXiv:2605.07846v1 Announce Type: new Abstract: Coarse-mask local image editing asks a model to modify a user-indicated region while preserving the surrounding scene. In practice, however, rough masks

Bridging the Last Mile of Circuit Design: PostEDA-Bench, a Hierarchical Benchmark for PPA Convergence and DRC Fixing

Model ReleasesDGX agent

arXiv:2605.06936v1 Announce Type: cross Abstract: LLM-based agents are increasingly applied to the 'last mile' of Electronic Design Automation (EDA): repairing residual sign-off Design Rule Check (DRC

CA-SQL: Complexity-Aware Inference Time Reasoning for Text-to-SQL via Exploration and Compute Budget Allocation

Model ReleasesDGX agent

arXiv:2605.08057v1 Announce Type: cross Abstract: While recent advancements in inference-time learning have improved LLM reasoning on Text-to-SQL tasks, current solutions still struggle to perform wel

Can Agents Price a Reaction? Evaluating LLMs on Chemical Cost Reasoning

Model ReleasesDGX agent

arXiv:2605.07251v1 Announce Type: new Abstract: Large Language Models (LLMs) have become increasingly capable as tool-using agents, with benchmarks spanning diverse general agentic tasks. Yet rigorous

Can You Break RLVER? Probing Adversarial Robustness of RL-Trained Empathetic Agents

Model ReleasesDGX agent

arXiv:2605.07138v1 Announce Type: new Abstract: Reinforcement learning from verifiable emotion rewards RLVER has produced language models with strong empathetic performance, evaluated on benchmarks th

CarCrashNet: A Large-Scale Dataset and Hierarchical Neural Solver for Data-Driven Structural Crash Simulation

Model ReleasesDGX agent

arXiv:2605.07098v1 Announce Type: new Abstract: Crash simulation is a cornerstone of modern vehicle development because it reduces the need for costly physical prototypes, accelerates safety-driven de

Chain-based Distillation for Effective Initialization of Variable-Sized Small Language Models

Model ReleasesDGX agent

arXiv:2605.07783v1 Announce Type: new Abstract: Large language models (LLMs) achieve strong performance but remain costly to deploy in resource-constrained settings. Training small language models (SL

ChartREG++: Towards Benchmarking and Improving Chart Referring Expression Grounding under Diverse referring clues and Multi-Target Referring

Model ReleasesDGX agent

arXiv:2605.07415v1 Announce Type: cross Abstract: Referring expression grounding is a core problem in visual grounding and is widely used as a diagnostic of spatial grounding and reasoning in vision a

Clinically Aware Synthetic Image Generation for Concept Coverage in Chest X-ray Models

Model ReleasesDGX agent

arXiv:2603.15525v2 Announce Type: replace Abstract: Deep learning models for chest X-ray diagnosis are constrained by limited coverage of clinically meaningful concept combinations in publicly availab

Cloud Storage Rapid: Turbocharged object storage for AI and analytics

Model ReleasesDGX agent

At Google Cloud Next ’26 we announced Cloud Storage Rapid, a family of object storage capabilities for data-intensive workloads like AI and analytics. Out of the gate, Cloud Storage Rapid consists of

Cluster-level reliability for trillion-parameter models on TPUs

Model ReleasesDGX agent

Frontier AI models have redefined the unit of compute. At trillion-parameter scale, AI training requires thousands of interconnected components, orchestrated in industrial-scale deployments to operate

CoCoReviewBench: A Completeness- and Correctness-Oriented Benchmark for AI Reviewers

Model ReleasesDGX agent

arXiv:2605.07905v1 Announce Type: cross Abstract: Despite the rapid development of AI reviewers, evaluating such systems remains challenging: metrics favor overlap with human reviews over correctness.

CommandSwarm: Safety-Aware Natural Language-to-Behavior-Tree Generation for Robotic Swarms

Model ReleasesDGX agent

arXiv:2605.07764v1 Announce Type: new Abstract: Natural-language interfaces can make swarm robotics more accessible to non-expert operators, but they must translate ambiguous user intent into executab

Continually Evolving Skill Knowledge in Vision Language Action Model

Model ReleasesDGX agent

arXiv:2511.18085v4 Announce Type: replace-cross Abstract: Vision-language-action (VLA) models show promising knowledge accumulation ability from pretraining, yet continual learning in VLA remains chal

Contrast-X: A Multi-Modal Contrast Image Synthesis Benchmark and Universal Modality Flow Matching

Model ReleasesDGX agent

arXiv:2601.15884v2 Announce Type: replace Abstract: Contrast-enhanced imaging is central to oncologic diagnosis, but contrast agents can be contraindicated for many of the patients who need them most.

Convergence and Emergence of In-Context Reinforcement Learning with Chain of Thought

Model ReleasesDGX agent

arXiv:2605.07123v1 Announce Type: new Abstract: In-context reinforcement learning (ICRL) refers to the ability of RL agents to adapt to new tasks at inference time without parameter updates by conditi

Convex Optimization with Nested Evolving Feasible Sets

Model ReleasesDGX agent

arXiv:2605.07386v1 Announce Type: new Abstract: Convex Optimization with Nested Evolving Feasible Sets (CONES)} is considered where the objective function f remains fixed but the feasible region evolv

CrossCult-KIBench: A Benchmark for Cross-Cultural Knowledge Insertion in MLLMs

Model ReleasesDGX agent

arXiv:2605.06115v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs), trained primarily on English-centric data, frequently generate culturally inappropriate or misaligned resp

CSMCIR: CoT-Enhanced Symmetric Alignment with Memory Bank for Composed Image Retrieval

Model ReleasesDGX agent

arXiv:2601.03728v3 Announce Type: replace-cross Abstract: Composed Image Retrieval (CIR) enables users to search for target images using both a reference image and manipulation text, offering substant

CSR: Infinite-Horizon Real-Time Policies with Massive Cached State Representations

Model ReleasesDGX agent

arXiv:2605.07325v1 Announce Type: cross Abstract: Deploying massive large language models (LLMs) as continuous cognitive engines for robotics is bottlenecked by the time-to-first-token (TTFT) latency

Curvature Beyond Positivity: Greedy Guarantees for Arbitrary Submodular Functions

Model ReleasesDGX agent

arXiv:2605.07902v1 Announce Type: new Abstract: Submodular functions -- functions exhibiting diminishing returns -- are central to machine learning. When the objective is monotone and non-negative, th

CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios

Model ReleasesDGX agent

arXiv:2605.07830v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as autonomous agents in offensive cybersecurity. In this paper, we reveal an interesting phenom

Data Contamination in Neural Hieroglyphic Translation: A Reproducibility Study

Model ReleasesDGX agent

arXiv:2605.07453v1 Announce Type: new Abstract: Ancient and endangered languages pose a unique challenge for NLP: their datasets are inherently scarce, difficult to expand, and built from formulaic co

Dataset Watermarking for Closed LLMs with Provable Detection

Model ReleasesDGX agent

arXiv:2605.06865v1 Announce Type: new Abstract: Large language models (LLMs) are pre-trained and post-trained on vast amounts of loosely curated data, raising the possibility that these models may hav

Decoding the Pulse of Reasoning VLMs in Multi-Image Understanding Tasks

Model ReleasesDGX agent

arXiv:2603.04676v2 Announce Type: replace-cross Abstract: Multi-image reasoning remains a significant challenge for vision-language models (VLMs). We investigate a previously overlooked phenomenon: du

Deeply Dual Supervised learning for melanoma recognition

Model ReleasesDGX agent

arXiv:2508.01994v2 Announce Type: replace Abstract: As the application of deep learning in dermatology continues to grow, the recognition of melanoma has garnered significant attention, demonstrating

DeepSeek V4 Flash is ~90% cheaper than GPT 5.4 Mini and ~70% cheaper than Gemini 3.1 Flash Lite For devs pushing ~500M tok/month, this is th…

Model ReleasesDGX agent

DeepSeek V4 Flash is ~90% cheaper than GPT 5.4 Mini and ~70% cheaper than Gemini 3.1 Flash Lite For devs pushing ~500M tok/month, this is the difference between: GPT 5.4 Mini: ~394/mo Gemini 3.1 Flash

DeepSeek V4 Pro brings long-context reasoning and SOTA coding performance to Together AI serverless. The next layer is serving it efficientl…

Model ReleasesDGX agent

DeepSeek V4 Pro brings long-context reasoning and SOTA coding performance to Together AI serverless. The next layer is serving it efficiently: KV cache, prefix reuse, hybrid attention, batching, kerne

Delulu: A Verified Multi-Lingual Benchmark for Code Hallucination Detection in Fill-in-the-Middle Tasks

Model ReleasesDGX agent

arXiv:2605.07024v1 Announce Type: new Abstract: Large Language Models for code generation frequently produce hallucinations in Fill-in-the-Middle (FIM) tasks -- plausible but incorrect completions suc

Detecting Distillation Data from Reasoning Models

Model ReleasesDGX agent

arXiv:2510.04850v3 Announce Type: replace-cross Abstract: Reasoning distillation has emerged as a prevailing paradigm for transferring reasoning capabilities from large reasoning models to small langu

Dino U-Net: Exploiting High-Fidelity Dense Features from Foundation Models for Medical Image Segmentation

Model ReleasesDGX agent

arXiv:2508.20909v2 Announce Type: replace Abstract: Foundation models pre-trained on large-scale natural image datasets offer a powerful paradigm for medical image segmentation. However, effectively t

Disagreement-Regularized Importance Sampling for Adversarial Label Corruption

Model ReleasesDGX agent

arXiv:2605.07551v1 Announce Type: new Abstract: Standard Importance Sampling (IS) collapses under label corruption because high-norm examples, prioritized for variance reduction, are often adversarial

Discord launches Nitro Rewards, giving Nitro subscribers access to offers from gaming services like Xbox Game Pass and hardware like Logitech G at no extra cost (Amanda Silberling/TechCrunch)

Model ReleasesDGX agent

Amanda Silberling / TechCrunch: Discord launches Nitro Rewards, giving Nitro subscribers access to offers from gaming services like Xbox Game Pass and hardware like Logitech G at no extra cost — Disco

Discovering Ordinary Differential Equations with LLM-Based Qualitative and Quantitative Evaluation

Model ReleasesDGX agent

arXiv:2605.07323v1 Announce Type: new Abstract: Discovering governing differential equations from observational data is a fundamental challenge in scientific machine learning. Existing symbolic regres

Divide and Conquer: Object Co-occurrence Helps Mitigate Simplicity Bias in OOD Detection

Model ReleasesDGX agent

arXiv:2605.07821v1 Announce Type: cross Abstract: Out-of-distribution (OOD) detection is crucial for ensuring the reliability of deep learning models. Existing methods mostly focus on regular entangle

DKDS: A Benchmark Dataset of Degraded Kuzushiji Documents with Seals for Detection and Binarization

Model ReleasesDGX agent

arXiv:2511.09117v4 Announce Type: replace Abstract: Kuzushiji, a pre-modern Japanese cursive script, can currently be read and understood by only a few thousand trained experts in Japan. With the rapi

Do Joint Audio-Video Generation Models Understand Physics?

Model ReleasesDGX agent

arXiv:2605.07061v1 Announce Type: cross Abstract: Joint audio-video generation models are rapidly approaching professional production quality, raising a central question: do they understand audio-visu

Domain-level metacognitive monitoring in frontier LLMs: A 33-model atlas

Model ReleasesDGX agent

arXiv:2605.06673v1 Announce Type: cross Abstract: Aggregate metacognitive quality scores mask within-model variation across MMLU benchmark domains. We administered 1,500 MMLU items (250 per domain, un

DRIP-R: A Benchmark for Decision-Making and Reasoning Under Real-World Policy Ambiguity in the Retail Domain

Model ReleasesDGX agent

arXiv:2605.07699v1 Announce Type: cross Abstract: LLM-based agents are increasingly deployed for routine but consequential tasks in real-world domains, where their behavior is governed by inherently a

DT-PBO: an Interpretable Tree-based Surrogate Model for Preferential Bayesian Optimization

Model ReleasesDGX agent

arXiv:2512.14263v2 Announce Type: replace-cross Abstract: Preferential Bayesian Optimization (PBO) aims to find a decision-maker's most preferred solution in as few pairwise comparisons as possible. E

Echo: KV-Cache-Free Associative Recall with Spectral Koopman Operators

Model ReleasesDGX agent

arXiv:2605.06997v1 Announce Type: new Abstract: Long chain-of-thought reasoning and agentic tool-calling produce traces spanning tens of thousands of tokens, yet Transformer KV caches grow linearly wi

Effective and Memory-Efficient Alternatives to ECC for Reliable Large-Scale DNNs

Model ReleasesDGX agent

arXiv:2605.07417v1 Announce Type: cross Abstract: Modern Deep Learning (DL) workloads are increasingly deployed in safety-critical domains, such as automotive systems and hyperscale data centers, wher

Efficient Data Selection for Multimodal Models via Incremental Optimization Utility

Model ReleasesDGX agent

arXiv:2605.07488v1 Announce Type: new Abstract: The scaling of Large Multimodal Models (LMMs) is constrained by the quality-quantity trade-off inherent in synthetic data. Previous approaches, such as

EgoPro-Bench: Benchmarking Personalized Proactive Interaction in Egocentric Video Streams

Model ReleasesDGX agent

arXiv:2605.07299v1 Announce Type: cross Abstract: Existing Multimodal Large Language Models (MLLMs) remain primarily reactive, failing to continuously perceive environments or proactively assist users

End-to-end PDDL Planning with Hardcoded and Dynamic Agents

Model ReleasesDGX agent

arXiv:2512.09629v2 Announce Type: replace Abstract: We present an end-to-end framework for planning supported by verifiers. An orchestrator receives a human specification written in natural language a

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation

Model ReleasesDGX agent

arXiv:2605.07247v1 Announce Type: new Abstract: Scalable AI agents training relies on interactive environments that faithfully simulate the consequences of agent actions. Manually crafted environments

ESSAM: A Novel Competitive Evolution Strategies Approach to Reinforcement Learning for Memory Efficient LLMs Fine-Tuning

Model ReleasesDGX agent

arXiv:2602.01003v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has become a key training step for improving mathematical reasoning in large language models (LLMs), but it often

Evaluating Large Language Models in Scientific Discovery

Model ReleasesDGX agent

arXiv:2512.15567v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly applied to scientific research, yet prevailing science benchmarks probe decontextualized knowledge and

Evaluating Prompt Injection Defenses for Educational LLM Tutors: Security-Usability-Latency Trade-offs

Model ReleasesDGX agent

arXiv:2605.06669v1 Announce Type: cross Abstract: Educational LLM tutors face a core AI alignment challenge: they must follow user intent while preserving pedagogical constraints and safety policies.

Ever wished your agent could read PDFs, images, and Office documents as easily as plain text? Or combine the safety of a secure sandbox with…

Model ReleasesDGX agent

Ever wished your agent could read PDFs, images, and Office documents as easily as plain text? Or combine the safety of a secure sandbox with the full power of Bash access? We built exactly that. Meet

Exact Flow Linear Attention: Exact Solution from Continuous-Time Dynamics

Model ReleasesDGX agent

arXiv:2512.12602v4 Announce Type: replace Abstract: In this paper, we introduce Exact Flow Linear Attention~(EFLA), an exact-flow formulation of delta-rule linear attention. We show that the delta-rul

Exact Is Easier: Credit Assignment for Cooperative LLM Agents

Model ReleasesDGX agent

arXiv:2603.06859v2 Announce Type: replace-cross Abstract: Removing an agent from a cooperative team to measure its contribution seems natural, yet in multi-agent LLM systems this evaluation distorts t

Excluding the Target Domain Improves Extrapolation: Deconfounded Hierarchical Physics Constraints

Model ReleasesDGX agent

arXiv:2605.07485v1 Announce Type: cross Abstract: Extrapolation to out-of-distribution conditions is a fundamental challenge for physics-constrained deep generative models. Existing methods apply phys

← Previous
1…271272273274275…377
Next →