AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,606Total entries
1Added by human
84,605Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,593 results
20 May 2026

Hallucination as Exploit: Evidence-Carrying Multimodal Agents

Model ReleasesDGX agent

arXiv:2605.19192v1 Announce Type: new Abstract: Multimodal agents use screenshots, documents, and webpages to choose tool calls. When a false visual claim triggers a click, email, extraction, or trans

HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models

Model ReleasesDGX agent

arXiv:2605.19341v1 Announce Type: cross Abstract: Hallucination remains a central failure mode of large language models, but existing benchmarks operationalize it inconsistently across summarization,

HAVEN: Hierarchically Aligned Multimodal Benchmark for Unified Video Understanding

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.19223v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) exhibit strong performance on standard video tasks, their ability to faithfully summarize and reason over

HELLoRA: Hot Experts Layer-Level Low-Rank Adaptation for Mixture-of-Experts Models

Model ReleasesDGX agent

arXiv:2605.18795v1 Announce Type: cross Abstract: Low-Rank Adaptation (LoRA) dominates parameter-efficient fine-tuning of large language models, yet most variants target dense architectures. Mixture-o

How do LLMs Compute Verbal Confidence

Model ReleasesDGX agent

arXiv:2603.17839v3 Announce Type: replace-cross Abstract: Verbal confidence -- prompting LLMs to state their confidence as a number or category -- is widely used to extract uncertainty estimates from

How Faithful Is Trajectory-Based Data Attribution? Error Sources, Remedies, and Practical Guidelines

Model ReleasesDGX agent

arXiv:2605.18814v1 Announce Type: new Abstract: Trajectory-based data attribution methods estimate the influence of training samples on model predictions by unrolling the training trajectory. They are

How Far Are We From True Auto-Research?

Model ReleasesDGX agent

arXiv:2605.19156v1 Announce Type: new Abstract: Recent auto-research systems can produce complete papers, but feasibility is not the same as quality, and the field still lacks a systematic study of ho

How Ramp engineers accelerate code review with Codex

Model ReleasesDGX agent

Ramp engineers use OpenAI's Codex to streamline and accelerate their code review process, leveraging AI-assisted code understanding and analysis. The implementation likely demonstrates how Codex helps

Hybrid-LoRA: Bridging Full Fine-Tuning and Low-Rank Adaptation for Post-Training

Model ReleasesDGX agent

arXiv:2605.18822v1 Announce Type: cross Abstract: Post-training has become essential for adapting large language models (LLMs) to complex downstream behaviors, including instruction following, prefere

I am starting to have trouble paying attention to even interesting information if it is written in Claude or ChatGPT house style. I think so…

Model ReleasesDGX agent

I am starting to have trouble paying attention to even interesting information if it is written in Claude or ChatGPT house style. I think some is the sameness of the rhythm rather than obvious tics: C

I don't have much to say about this year's Google I/O because I prefer to write about products that have shipped, not just 'coming soon' ann…

Model ReleasesDGX agent

I don't have much to say about this year's Google I/O because I prefer to write about products that have shipped, not just 'coming soon' announcements - but here are some notes on Gemini Spark and Ant

i don't think i need cloud models anymore

Model ReleasesDGX agent

i don't think i need cloud models anymore MTP speedup Qwen by 2.5x in Atomic Chat Dense vs MoE models on 2x RTX 5090 Qwen3.6 27B: 51 → 117 tps +137% Qwen3.6 35B-A3B: 218 → 267 tps +25% MTP drafts seve

iGSP:Implicit Gradient Subspace Projection for Efficient Continual Learning of Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.19301v1 Announce Type: new Abstract: Vision-Language Models require efficient adaptation to continually emerging downstream tasks. While Parameter-Efficient Fine-Tuning mitigates catastroph

IMLJD: A Computational Dataset for Indian Matrimonial Litigation Analysis

Model ReleasesDGX agent

arXiv:2605.19346v1 Announce Type: cross Abstract: We present IMLJD, an open dataset of 3,613 Indian court judgments covering matrimonial disputes under IPC Section 498A, the Protection of Women from D

In-Context Learning Operates as Concept Subspace Learning

Model ReleasesDGX agent

arXiv:2605.18830v1 Announce Type: new Abstract: Regression and Bayesian accounts of in-context learning (ICL) explain how demonstrations can induce predictors, while mechanistic analyses often identif

Information Processing Capacity of Stationary Physical Systems: Theory, Data-efficient Estimation Methods, and Photonic Demonstration

Model ReleasesDGX agent

arXiv:2605.19152v1 Announce Type: cross Abstract: Physical computing systems provide a promising route toward hardware-native machine learning, but their computational capabilities remain difficult to

INSHAPE: Instance-Level Shapelets for Interpretable Time-Series Classification

Model ReleasesDGX agent

arXiv:2605.20088v1 Announce Type: cross Abstract: Discovering shapelets -- i.e., discriminative temporal patterns within time series -- has been widely studied to address the inherent complexity of ti

Introducing Agent Executor, Google’s distributed Agent Runtime

Model ReleasesDGX agent

As models and harnesses improve, agents are taking on increasingly complex tasks that can run for hours or even days. But as we push agents to do more, this has surfaced a new operational problem: lon

Introducing: Cohere Command A+ We’ve created our most powerful LLM yet, optimized it to run on as little hardware as possible, and released …

Model ReleasesDGX agent

Cohere announced Command A+, their most advanced large language model to date, designed with optimization for efficient hardware requirements to enable broader deployment and accessibility. The model

It's been *almost* a bit quiet around LLM architecture releases in the past two weeks 😅 Interesting tidbit is the parallel block design. Vi…

Model ReleasesDGX agent

It's been *almost* a bit quiet around LLM architecture releases in the past two weeks 😅 Interesting tidbit is the parallel block design. Via the Cmd-A the tech report 'equivalent performance but signi

it’s in gemini, just create it in ai studio. oh, that’s for your personal google one account. for workspace you need gemini business. no, no…

Model ReleasesDGX agent

it’s in gemini, just create it in ai studio. oh, that’s for your personal google one account. for workspace you need gemini business. no, not gemini advanced, that’s ai pro now. unless you need ai ult

JAXenstein: Accelerated Benchmarking for First-Person Environments

Model ReleasesDGX agent

arXiv:2605.19926v1 Announce Type: new Abstract: The progression of reinforcement learning algorithms have been driven by challenging benchmarks. The rate in which a researcher can iterate on a problem

K-Quantization and its Impact on Output Performance

Model ReleasesDGX agent

arXiv:2605.19645v1 Announce Type: new Abstract: Recent advancements in large language models (LLMs) have shown their remarkable capacities in many NLP tasks. However, their substantial size often pres

KappaPlace: Learning Hyperspherical Uncertainty for Visual Place Recognition via Prototype-Anchored Supervision

Model ReleasesDGX agent

arXiv:2605.19435v1 Announce Type: cross Abstract: Visual Place Recognition (VPR) is critical for autonomous navigation, yet state-of-the-art methods lack well-calibrated uncertainty estimation. Standa

Learn-by-Wire Training Control Governance: Bounded Autonomous Training Under Stress for Stability and Efficiency

Model ReleasesDGX agent

arXiv:2605.19008v1 Announce Type: new Abstract: Modern language-model training is increasingly exposed to instability, degraded runs, and wasted compute, especially under aggressive learning-rate, sca

Learning-Accelerated Optimization-based Trajectory Planning for Cooperative Aerial-Ground Handover Missions

Model ReleasesDGX agent

arXiv:2605.19562v1 Announce Type: cross Abstract: This paper presents a learning-augmented trajectory planning framework for cooperative unmanned aerial vehicle (UAV) and unmanned ground vehicle (UGV)

Learning Efficient Guardrails for Compliance

Model ReleasesDGX agent

arXiv:2510.03485v2 Announce Type: replace Abstract: Autonomous web agents are increasingly deployed for long-horizon tasks, yet their ability to adhere to real-world policies remains critically undere

Learning When to Adapt

Model ReleasesDGX agent

arXiv:2605.19028v1 Announce Type: new Abstract: Low-rank adaptation (LoRA) is a widely used parameter-efficient fine-tuning method, yet its learned correction is static: the same low-rank update is ap

Lens Privacy Sealing: A New Benchmark and Method for Physical Privacy-Preserving Action Recognition

Model ReleasesDGX agent

arXiv:2605.19578v1 Announce Type: cross Abstract: RGB camera-based surveillance systems enable human action recognition for public safety and healthcare, yet raise serious privacy concerns. Existing m

Less Back-and-Forth: A Comparative Study of Structured Prompting

Model ReleasesDGX agent

arXiv:2605.20149v1 Announce Type: cross Abstract: Large language models (LLMs) are widely used for open-ended tasks, but underspecified prompts can lead to low-quality answers and additional interacti

Library Hallucinations in LLM-Generated Code: A Risk Analysis Grounded in Developer Queries

Model ReleasesDGX agent

arXiv:2509.22202v3 Announce Type: replace-cross Abstract: Large language models (LLMs) now play a central role in code generation, yet they continue to hallucinate, frequently inventing non-existent l

LIFT and PLACE: A Simple, Stable, and Effective Knowledge Distillation Framework for Lightweight Diffusion Models

Model ReleasesDGX agent

arXiv:2605.19729v1 Announce Type: cross Abstract: We demonstrate that in knowledge distillation for diffusion models, the teacher network's highly complex denoising process - stemming from its substan

Lightweight and Fast Backdoor Model Detection

Model ReleasesDGX agent

arXiv:2605.18907v1 Announce Type: cross Abstract: Deep neural networks (DNN), despite their remarkable performance, are highly vulnerable to backdoor attacks. Existing defenses mainly rely on activati

LLM Benchmark Datasets Should Be Contamination-Resistant

Model ReleasesDGX agent

arXiv:2605.19999v1 Announce Type: cross Abstract: Benchmark datasets are critical for reproducible, reliable, and discriminative evaluation of LLMs. However, recent studies reveal that many benchmark

LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening

Model ReleasesDGX agent

arXiv:2605.19597v1 Announce Type: new Abstract: Evaluating large language models (LLMs) on natural-language logical reasoning is essential because rule-governed tasks require conclusions to follow str

LMM-Track4D: Eliciting 4D Dynamic Reasoning in LMMs via Trajectory-Grounded Dialogue

Model ReleasesDGX agent

arXiv:2605.19390v1 Announce Type: new Abstract: Recent large multimodal models (LMMs) have become increasingly capable on image and video understanding, yet still struggle to sustain 4D continuous spa

LWiAI Podcast #245 - TML-Interaction, Claude For Legal, Sam Altman on Stand

Model ReleasesDGX agent

This podcast episode covers three main topics: TML-Interaction (likely a new AI model or technical development), the application of Claude AI in legal settings and use cases, and Sam Altman's testimon

Lying Is Just a Phase: The Hidden Alignment Transition in Language Model Scaling

Model ReleasesDGX agent

arXiv:2605.18838v1 Announce Type: cross Abstract: Scaling laws predict loss from compute but not how capabilities interact. We measure the coupling between reasoning and truthfulness across 63 base mo

Lynx: Enabling Efficient MoE Inference through Dynamic Batch-Aware Expert Selection

Model ReleasesDGX agent

arXiv:2411.08982v3 Announce Type: replace Abstract: Selective parameter activation provided by Mixture-of-Expert (MoE) models have made them a popular choice in modern foundational models. However, Mo

m3BERT: A Modern, Multi-lingual, Matryoshka Bidirectional Encoder

Model ReleasesDGX agent

arXiv:2605.19568v1 Announce Type: new Abstract: Embedding models are pivotal in industrial information retrieval systems like search and advertising. However, existing pretrained models often exhibit

MAM-CLIP: Vision-Language Pretraining on Mammography Atlases for BI-RADS Classification

Model ReleasesDGX agent

arXiv:2605.19359v1 Announce Type: new Abstract: Deep learning methods have demonstrated promising results in predicting BI-RADS scores from mammography images. However, the interpretation of these ima

Managed Agents through the Gemini API is @GoogleAI's response to Anthropic Managed Agents Since it's powered by the new Antigravity agent bu…

Model ReleasesDGX agent

Managed Agents through the Gemini API is @GoogleAI's response to Anthropic Managed Agents Since it's powered by the new Antigravity agent built on Gemini 3.5 Flash, it is the most cost-effective gener

MANGO: Meta-Adaptive Network Gradient Optimization for Online Continual Learning

Model ReleasesDGX agent

arXiv:2605.19080v1 Announce Type: cross Abstract: In Online Continual Learning (OCL), a neural network sequentially learns from a non-stationary data stream in a single-pass with access only to a limi

Mask-to-Correct^+: Leveraging Retriever Diversity for Masking-guided Faithful Fact Correction

Model ReleasesDGX agent

arXiv:2605.18776v1 Announce Type: cross Abstract: The rapid spread of misinformation on social media highlights the need for robust, automated fact correction frameworks. However, existing works rely

Mathematical Reasoning in Large Language Models: Benchmarks, Architectures, Evaluation, and Open Challenges

Model ReleasesDGX agent

arXiv:2605.19723v1 Announce Type: cross Abstract: Mathematical reasoning is essential for problem-solving in education, science, and industry, serving as a crucial benchmark for evaluating artificial

Measuring Safety Alignment Effects in Autonomous Security Agents

Model ReleasesDGX agent

arXiv:2605.19722v1 Announce Type: cross Abstract: Do stock safety-aligned language models and their uncensored or abliterated derivatives behave differently when run as autonomous security agents? Sin

MedFM-Robust: Benchmarking Robustness of Medical Foundation Models

Model ReleasesDGX agent

arXiv:2605.19027v1 Announce Type: new Abstract: Medical foundation models (MedFMs) have emerged as transformative tools in healthcare, demonstrating capabilities across diverse clinical applications.

Meet Stable Audio 3.0, the model family built for artistic experimentation with open-weight models

Model ReleasesDGX agent

Stable Audio 3.0 is Stability AI's latest generative audio model designed to create music and soundscapes from textual prompts. The release includes four models ranging from 459M to 2.7B parameters, w

// Memory as a Model // The paper augments any LLM with a separate trained memory model that stores, retrieves, and integrates facts on its …

Model ReleasesDGX agent

// Memory as a Model // The paper augments any LLM with a separate trained memory model that stores, retrieves, and integrates facts on its behalf. It decouples memory updates from base-model weight u

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems

Model ReleasesDGX agent

arXiv:2605.19307v1 Announce Type: new Abstract: Visual Question Answering (VQA), as the representative multimodal task, serves as a key benchmark for evaluating the reasoning capabilities of Multimoda

Microsoft Senior AI developer just showed how they build AI agents with Claude at Microsoft. 34-minutes. free. By Microsoft team Opus 4.7 + …

Model ReleasesDGX agent

Microsoft Senior AI developer just showed how they build AI agents with Claude at Microsoft. 34-minutes. free. By Microsoft team Opus 4.7 + 1,400+ pre-built MCP tools plug Claude into agent → give it

MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency

Model ReleasesDGX agent

arXiv:2510.25897v2 Announce Type: replace Abstract: The default paradigm of post-training text-to-image generators includes post-hoc selection of generated images, and subsequent training with one rew

MixRea: Benchmarking Explicit-Implicit Reasoning in Large Language Models

Model ReleasesDGX agent

arXiv:2605.20128v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly integrated into high-stakes decision-making. Inspired by the theory of inattentional blindness in human co

MotionMERGE: A Multi-granular Framework for Human Motion Editing, Reasoning, Generation, and Explanation

Model ReleasesDGX agent

arXiv:2605.18956v1 Announce Type: new Abstract: Recent motion-language models unify tasks like comprehension and generation but operate at a coarse granularity, lacking fine-grained understanding and

MSAlign: Aligning Molecule and Mass Spectra Foundation Models for Metabolite Identification

Model ReleasesDGX agent

arXiv:2605.19752v1 Announce Type: new Abstract: Accurately identifying metabolites i.e. small molecules from mass spectrometry data remains a core challenge in metabolomics, with broad applications in

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation

Model ReleasesDGX agent

arXiv:2605.20183v1 Announce Type: new Abstract: Video generation is rapidly evolving from single-shot synthesis to complex multi-shot audio-video (MSAV) narratives to meet real-world demands. However,

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training

Model ReleasesDGX agent

arXiv:2510.18830v2 Announce Type: replace Abstract: The adoption of long context windows has become a standard feature in Large Language Models (LLMs), as extended contexts significantly enhance their

Multi-axis Analysis of Image Manipulation Localization

Model ReleasesDGX agent

arXiv:2605.20174v1 Announce Type: new Abstract: Advanced image editing software enables easy creation of highly convincing image manipulations, which has been made even more accessible in recent years

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs

Model ReleasesDGX agent

arXiv:2511.14159v2 Announce Type: replace Abstract: Evaluating the robustness of Large Vision-Language Models (LVLMs) is essential for their continued development and responsible deployment in real-wo

Next-Acceleration-Scale Prediction for Autoregressive MRI Reconstruction

Model ReleasesDGX agent

arXiv:2605.19354v1 Announce Type: cross Abstract: MRI reconstruction is an inherently ill-posed inverse problem, since incomplete measurements admit many plausible solutions. This ambiguity becomes mo

← Previous
1…229230231232233…377
Next →