AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,292 results
17 Apr 2026

Knowing When Not to Answer: Evaluating Abstention in Multimodal Reasoning Systems

Model ReleasesDGX agent

arXiv:2604.14799v1 Announce Type: new Abstract: Effective abstention (EA), recognizing evidence insufficiency and refraining from answering, is critical for reliable multimodal systems. Yet existing e

Label-efficient underwater species classification with logistic regression on frozen foundation model embeddings

Model ReleasesDGX agent

arXiv:2604.00313v2 Announce Type: replace Abstract: Automated species classification from underwater imagery is bottlenecked by the cost of expert annotation, and supervised models trained on one data

Last week, Anthropic announced Project Glasswing alongside Claude Mythos Preview, a model they described as so powerful at finding vulnerabi…


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases
DGX agent

Last week, Anthropic announced Project Glasswing alongside Claude Mythos Preview, a model they described as so powerful at finding vulnerabilities they couldn't release it. The announcement featured A

Learning Where to Embed: Noise-Aware Positional Embedding for Query Retrieval in Small-Object Detection

Model ReleasesDGX agent

arXiv:2604.15065v1 Announce Type: new Abstract: Transformer-based detectors have advanced small-object detection, but they often remain inefficient and vulnerable to background-induced query noise, wh

Let's talk content faithfulness. Four days ago, we launched ParseBench, the first document OCR benchmark for AI agents. Its most fundamental…

Model ReleasesDGX agent

Let's talk content faithfulness. Four days ago, we launched ParseBench, the first document OCR benchmark for AI agents. Its most fundamental metric asks: did the parser capture all the text, in order,

LexGenius: An Expert-Level Benchmark for Large Language Models in Legal General Intelligence

Model ReleasesDGX agent

arXiv:2512.04578v3 Announce Type: replace Abstract: Legal general intelligence (GI) refers to artificial intelligence (AI) that encompasses legal understanding, reasoning, and decision-making, simulat

Listen, Correct, and Feed Back: Spoken Pedagogical Feedback Generation

Model ReleasesDGX agent

arXiv:2604.14177v1 Announce Type: new Abstract: Grammatical error correction (GEC) and explanation (GEE) have made rapid progress, but real teaching scenarios also require learner-friendly pedagogical

Listen to the OpenAI Podcast on— Spotify https://open.spotify.com/show/0zojMEDizKMh3aTxnGLENP Apple https://podcasts.apple.com/us/podcast/op…

Model ReleasesDGX agent

Listen to the OpenAI Podcast on— Spotify https://open.spotify.com/show/0zojMEDizKMh3aTxnGLENP Apple https://podcasts.apple.com/us/podcast/openai-podcast/id1820330260 YouTube https://youtu.be/UZyH0nx5z

LLM Predictive Scoring and Validation: Inferring Experience Ratings from Unstructured Text

Model ReleasesDGX agent

arXiv:2604.14321v1 Announce Type: new Abstract: We tasked GPT-4.1 to read what baseball fans wrote about their game-day experience and predict the overall experience rating each fan gave on a 0-10 sur

LLMOrbit: A Circular Taxonomy of Large Language Models -From Scaling Walls to Agentic AI Systems

Model ReleasesDGX agent

arXiv:2601.14053v2 Announce Type: replace-cross Abstract: The field of artificial intelligence has undergone a revolution from foundational Transformer architectures to reasoning-capable systems appro

LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking

Model ReleasesDGX agent

arXiv:2604.15149v1 Announce Type: new Abstract: As reinforcement Learning with Verifiable Rewards (RLVR) has become the dominant paradigm for scaling reasoning capabilities in LLMs, a new failure mode

LongAct: Harnessing Intrinsic Activation Patterns for Long-Context Reinforcement Learning

Model ReleasesDGX agent

arXiv:2604.14922v1 Announce Type: cross Abstract: Reinforcement Learning (RL) has emerged as a critical driver for enhancing the reasoning capabilities of Large Language Models (LLMs). While recent ad

MADE: A Living Benchmark for Multi-Label Text Classification with Uncertainty Quantification of Medical Device Adverse Events

Model ReleasesDGX agent

arXiv:2604.15203v1 Announce Type: new Abstract: Machine learning in high-stakes domains such as healthcare requires not only strong predictive performance but also reliable uncertainty quantification

Magnitude Is All You Need? Rethinking Phase in Quantum Encoding of Complex SAR Data

Model ReleasesDGX agent

arXiv:2604.14229v1 Announce Type: cross Abstract: Synthetic Aperture Radar (SAR) data is inherently complex-valued, while quantum machine learning (QML) models naturally operate in complex Hilbert spa

MARCA: A Checklist-Based Benchmark for Multilingual Web Search

Model ReleasesDGX agent

arXiv:2604.14448v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as sources of information, yet their reliability depends on the ability to search the web, select rel

MAS-Bench: A Unified Benchmark for Shortcut-Augmented Hybrid Mobile GUI Agents

Model ReleasesDGX agent

arXiv:2509.06477v2 Announce Type: replace Abstract: Shortcuts such as APIs and deep-links have emerged as efficient complements to flexible GUI operations, fostering a promising hybrid paradigm for ML

Maximal Brain Damage Without Data or Optimization: Disrupting Neural Networks via Sign-Bit Flips

Model ReleasesDGX agent

arXiv:2502.07408v2 Announce Type: replace-cross Abstract: Deep Neural Networks (DNNs) can be catastrophically disrupted by flipping only a handful of parameter bits. We introduce Deep Neural Lesion (D

Measuring multi-calibration

Model ReleasesDGX agent

arXiv:2506.11251v2 Announce Type: replace-cross Abstract: A suitable scalar metric can help measure multi-calibration, defined as follows. When the expected values of observed responses are equal to c

MEBench: A Novel Benchmark for Understanding Mutual Exclusivity Bias in Vision-Language Models

Model ReleasesDGX agent

arXiv:2505.20122v2 Announce Type: replace Abstract: This paper introduces MEBench, a novel benchmark for evaluating mutual exclusivity (ME) bias, a cognitive phenomenon observed in children during wor

Mechanistic Decoding of Cognitive Constructs in LLMs

Model ReleasesDGX agent

arXiv:2604.14593v1 Announce Type: new Abstract: While Large Language Models (LLMs) demonstrate increasingly sophisticated affective capabilities, the internal mechanisms by which they process complex

MemGround: Long-Term Memory Evaluation Kit for Large Language Models in Gamified Scenarios

Model ReleasesDGX agent

arXiv:2604.14158v1 Announce Type: new Abstract: Current evaluations of long-term memory in LLMs are fundamentally static. By fixating on simple retrieval and short-context inference, they neglect the

MetaDent: Labeling Clinical Images for Vision-Language Models in Dentistry

Model ReleasesDGX agent

arXiv:2604.14866v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have demonstrated significant potential in medical image analysis, yet their application in intraoral photography remains

Mitigating LLM biases toward spurious social contexts using direct preference optimization

Model ReleasesDGX agent

arXiv:2604.02585v2 Announce Type: replace-cross Abstract: LLMs are increasingly used for high-stakes decision-making, yet their sensitivity to spurious contextual information can introduce harmful bia

MixAtlas: Uncertainty-aware Data Mixture Optimization for Multimodal LLM Midtraining

Model ReleasesDGX agent

arXiv:2604.14198v1 Announce Type: cross Abstract: Domain reweighting can improve sample efficiency and downstream generalization, but data-mixture optimization for multimodal midtraining remains large

MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation

Model ReleasesDGX agent

arXiv:2604.15309v1 Announce Type: cross Abstract: The rapid progress of Artificial Intelligence Generated Content (AIGC) tools enables images, videos, and visualizations to be created on demand for we

Modular Continual Learning via Zero-Leakage Reconstruction Routing and Autonomous Task Discovery

Model ReleasesDGX agent

arXiv:2604.14375v1 Announce Type: new Abstract: Catastrophic forgetting remains a primary hurdle in sequential task learning for artificial neural networks. We propose a silicon-native modular archite

ModuSeg: Decoupling Object Discovery and Semantic Retrieval for Training-Free Weakly Supervised Segmentation

Model ReleasesDGX agent

arXiv:2604.07021v2 Announce Type: replace Abstract: Weakly supervised semantic segmentation aims to achieve pixel-level predictions using image-level labels. Existing methods typically entangle semant

Multigrain-aware Semantic Prototype Scanning and Tri-Token Prompt Learning Embraced High-Order RWKV for Pan-Sharpening

Model ReleasesDGX agent

arXiv:2604.14622v1 Announce Type: new Abstract: In this work, we propose a Multigrain-aware Semantic Prototype Scanning paradigm for pan-sharpening, built upon a high-order RWKV architecture and a tri

Neuro-Oracle: A Trajectory-Aware Agentic RAG Framework for Interpretable Epilepsy Surgical Prognosis

Model ReleasesDGX agent

arXiv:2604.14216v1 Announce Type: cross Abstract: Predicting post-surgical seizure outcomes in pharmacoresistant epilepsy is a clinical challenge. Conventional deep-learning approaches operate on stat

OmniGCD: Abstracting Generalized Category Discovery for Modality Agnosticism

Model ReleasesDGX agent

arXiv:2604.14762v1 Announce Type: new Abstract: Generalized Category Discovery (GCD) challenges methods to identify known and novel classes using partially labeled data, mirroring human category learn

Open-Set Vein Biometric Recognition with Deep Metric Learning

Model ReleasesDGX agent

arXiv:2604.14874v1 Announce Type: new Abstract: Most state-of-the-art vein recognition methods rely on closed-set classification, which inherently limits their scalability and prevents the adaptive en

OpenAI ratchets up Codex’s agentic capabilities to rival Claude Code

Model ReleasesDGX agent

OpenAI Group PBC today announced a major revamp of its artificial intelligence coding tool Codex, giving it a number of new “agentic” capabilities that enable more complex task automation. The ChatGPT

OpenMobile: Building Open Mobile Agents with Task and Trajectory Synthesis

Model ReleasesDGX agent

arXiv:2604.15093v1 Announce Type: cross Abstract: Mobile agents powered by vision-language models have demonstrated impressive capabilities in automating mobile tasks, with recent leading models achie

Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games

Model ReleasesDGX agent

arXiv:2506.03610v3 Announce Type: replace Abstract: Large Language Model (LLM) agents are reshaping the game industry, by enabling more intelligent and human-preferable characters. Yet, current game b

Our virtual hackathon is back! Join us for a week of building with Opus 4.7 alongside developers from around the world. The Claude Code team…

Model ReleasesDGX agent

Our virtual hackathon is back! Join us for a week of building with Opus 4.7 alongside developers from around the world. The Claude Code team will be in the room all week, with a prize pool of $100K in

Pangu-ACE: Adaptive Cascaded Experts for Educational Response Generation on EduBench

Model ReleasesDGX agent

arXiv:2604.14828v1 Announce Type: new Abstract: Educational assistants should spend more computation only when the task needs it. This paper rewrites our earlier draft around the system that was actua

Parameter estimation for land-surface models using Neural Physics

Model ReleasesDGX agent

arXiv:2505.02979v3 Announce Type: replace-cross Abstract: We propose a novel inverse-modelling approach which estimates the parameters of a simple land-surface model (LSM) by assimilating data into a

PeerPrism: Peer Evaluation Expertise vs Review-writing AI

Model ReleasesDGX agent

arXiv:2604.14513v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used in scientific peer review, assisting with drafting, rewriting, expansion, and refinement. However, ex

Physical Intelligence says its new model, π0.7, can direct robots on tasks they weren't trained on, an 'early sign' of generalization, surprising researchers (Connie Loizos/TechCrunch)

Model ReleasesDGX agent

Connie Loizos / TechCrunch: Physical Intelligence says its new model, π0.7, can direct robots on tasks they weren't trained on, an “early sign” of generalization, surprising researchers — Physical Int

Physically-Induced Atmospheric Adversarial Perturbations: Enhancing Transferability and Robustness in Remote Sensing Image Classification

Model ReleasesDGX agent

arXiv:2604.14643v1 Announce Type: new Abstract: Adversarial attacks pose a severe threat to the reliability of deep learning models in remote sensing (RS) image classification. Most existing methods r

PolyBench: Benchmarking LLM Forecasting and Trading Capabilities on Live Prediction Market Data

Model ReleasesDGX agent

arXiv:2604.14199v1 Announce Type: cross Abstract: Predicting real-world events from live market signals demands systems that fuse qualitative news with quantitative order-book dynamics under strict te

POP: Prefill-Only Pruning for Efficient Large Model Inference

Model ReleasesDGX agent

arXiv:2602.03295v2 Announce Type: replace Abstract: Large Language Models (LLMs) and Vision-Language Models (VLMs) have demonstrated remarkable capabilities. However, their deployment is hindered by s

PortraitCraft: A Benchmark for Portrait Composition Understanding and Generation

Model ReleasesDGX agent

arXiv:2604.03611v2 Announce Type: replace Abstract: Portrait composition plays a central role in portrait aesthetics and visual communication, yet existing datasets and benchmarks mainly focus on coar

Prism: Symbolic Superoptimization of Tensor Programs

Model ReleasesDGX agent

arXiv:2604.15272v1 Announce Type: cross Abstract: This paper presents Prism, the first symbolic superoptimizer for tensor programs. The key idea is sGraph, a symbolic, hierarchical representation that

Prompt-Guided Image Editing with Masked Logit Nudging in Visual Autoregressive Models

Model ReleasesDGX agent

arXiv:2604.14591v1 Announce Type: new Abstract: We address the problem of prompt-guided image editing in visual autoregressive models. Given a source image and a target text prompt, we aim to modify t

Prompt Optimization Is a Coin Flip: Diagnosing When It Helps in Compound AI Systems

Model ReleasesDGX agent

arXiv:2604.14585v1 Announce Type: cross Abstract: Prompt optimization in compound AI systems is statistically indistinguishable from a coin flip: across 72 optimization runs on Claude Haiku (6 methods

ProRank: Prompt Warmup via Reinforcement Learning for Small Language Models Reranking

Model ReleasesDGX agent

arXiv:2506.03487v3 Announce Type: replace-cross Abstract: Reranking is fundamental to information retrieval and retrieval-augmented generation, with recent Large Language Models (LLMs) significantly a

Purging the Gray Zone: Latent-Geometric Denoising for Precise Knowledge Boundary Awareness

Model ReleasesDGX agent

arXiv:2604.14324v1 Announce Type: new Abstract: Large language models (LLMs) often exhibit hallucinations due to their inability to accurately perceive their own knowledge boundaries. Existing abstent

QuantCode-Bench: A Benchmark for Evaluating the Ability of Large Language Models to Generate Executable Algorithmic Trading Strategies

Model ReleasesDGX agent

arXiv:2604.15151v1 Announce Type: new Abstract: Large language models have demonstrated strong performance on general-purpose programming tasks, yet their ability to generate executable algorithmic tr

Query pipeline optimization for cancer patient question answering systems

Model ReleasesDGX agent

arXiv:2412.14751v2 Announce Type: replace Abstract: Retrieval-augmented generation (RAG) mitigates hallucination in Large Language Models (LLMs) by using query pipelines to retrieve relevant external

Regret Tail Characterization of Optimal Bandit Algorithms with Generic Rewards

Model ReleasesDGX agent

arXiv:2604.14876v1 Announce Type: cross Abstract: We study the tail behavior of regret in stochastic multi-armed bandits for algorithms that are asymptotically optimal in expectation. While minimizing

RELOAD: A Robust and Efficient Learned Query Optimizer for Database Systems

Model ReleasesDGX agent

arXiv:2604.14725v1 Announce Type: cross Abstract: Recent advances in query optimization have shifted from traditional rule-based and cost-based techniques towards machine learning-driven approaches. A

Rethinking Patient Education as Multi-turn Multi-modal Interaction

Model ReleasesDGX agent

arXiv:2604.14656v1 Announce Type: cross Abstract: Most medical multimodal benchmarks focus on static tasks such as image question answering, report generation, and plain-language rewriting. Patient ed

Retrieve, Then Classify: Corpus-Grounded Automation of Clinical Value Set Authoring

Model ReleasesDGX agent

arXiv:2604.14616v1 Announce Type: new Abstract: Clinical value set authoring -- the task of identifying all codes in a standardized vocabulary that define a clinical concept -- is a recurring bottlene

ReviewGrounder: Improving Review Substantiveness with Rubric-Guided, Tool-Integrated Agents

Model ReleasesDGX agent

arXiv:2604.14261v1 Announce Type: new Abstract: The rapid rise in AI conference submissions has driven increasing exploration of large language models (LLMs) for peer review support. However, LLM-base

Right at My Level: A Unified Multilingual Framework for Proficiency-Aware Text Simplification

Model ReleasesDGX agent

arXiv:2604.05302v2 Announce Type: replace Abstract: Text simplification supports second language (L2) learning by providing comprehensible input, consistent with the Input Hypothesis. However, constru

SafeHarness: Lifecycle-Integrated Security Architecture for LLM-based Agent Deployment

Model ReleasesDGX agent

arXiv:2604.13630v1 Announce Type: cross Abstract: The performance of large language model (LLM) agents depends critically on the execution harness, the system layer that orchestrates tool use, context

SAGE Celer 2.6 Technical Card

Model ReleasesDGX agent

arXiv:2604.14168v1 Announce Type: new Abstract: We introduce SAGE Celer 2.6, the latest in our line of general-purpose Celer models from SAGEA. Celer 2.6 is available in 5B, 10B, and 27B parameter siz

SAGE: Sign-Adaptive Gradient for Memory-Efficient LLM Optimization

Model ReleasesDGX agent

arXiv:2604.07663v2 Announce Type: replace Abstract: The AdamW optimizer, while standard for LLM pretraining, is a critical memory bottleneck, consuming optimizer states equivalent to twice the model's

SAQ: Stabilizer-Aware Quantum Error Correction Decoder

Model ReleasesDGX agent

arXiv:2512.08914v2 Announce Type: replace-cross Abstract: Quantum Error Correction (QEC) decoding faces a fundamental accuracy-efficiency tradeoff. Classical methods like Minimum Weight Perfect Matchi

← Previous
1…338339340341342…372
Next →