AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,606Total entries
1Added by human
84,605Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,593 results
19 May 2026

BacktestBench: Benchmarking Large Language Models for Automated Quantitative Strategy Backtesting

Model ReleasesDGX agent

arXiv:2605.17937v1 Announce Type: cross Abstract: Quantitative backtesting is essential for evaluating trading strategies but remains hampered by high technical barriers and limited scalability. While

Barriers for Learning in an Evolving World: Mathematical Understanding of Loss of Plasticity

Model ReleasesDGX agent

arXiv:2510.00304v3 Announce Type: replace-cross Abstract: Deep learning models excel in stationary data but struggle in non-stationary environments due to a phenomenon known as loss of plasticity (LoP

Bayesian-Monte Carlo Schedule Updating for Construction Digital Twins: A Probabilistic Framework for Dynamic Project Forecasting


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases
DGX agent

arXiv:2605.17608v1 Announce Type: cross Abstract: Construction projects frequently experience schedule delays and forecasting uncertainty due to variability in labor productivity, material availabilit

Beacon: Single-Turn Diagnosis and Mitigation of Latent Sycophancy in Large Language Models

Model ReleasesDGX agent

arXiv:2510.16727v2 Announce Type: replace-cross Abstract: Large language models internalize a structural trade-off between truthfulness and obsequious flattery, emerging from reward optimization that

Bench2Drive-Robust: Benchmarking Closed-Loop Autonomous Driving under Deployment Perturbations

Model ReleasesDGX agent

arXiv:2605.18059v1 Announce Type: new Abstract: Robustness is a critical requirement for deploying autonomous driving systems in the real world. Existing robustness benchmarks for autonomous driving h

Benchmarking inference at scale: coding agents

Model ReleasesDGX agent

This article presents benchmarking results for AI coding agents evaluated at scale, likely comparing performance metrics such as code generation accuracy, execution success rates, and inference effici

Benchmarking Mythos-Linked Bug Rediscovery

Model ReleasesDGX agent

arXiv:2605.17416v1 Announce Type: cross Abstract: Anthropic's April 2026 Mythos materials combine benchmark claims with concrete bug-finding stories across OpenBSD, FreeBSD, Linux, FFmpeg, and browser

Benchmarking Recurrent Event-Based Object Detection for Industrial Multi-Class Recognition on MTevent

Model ReleasesDGX agent

arXiv:2603.21787v2 Announce Type: replace Abstract: Event cameras are attractive for industrial robotics because they provide high temporal resolution, high dynamic range, and reduced motion blur. How

BESplit: Bias-Compensated Split Federated Learning with Evidential Aggregation

Model ReleasesDGX agent

arXiv:2605.17508v1 Announce Type: cross Abstract: Split Federated Learning (SFL) enables privacy-preserving collaborative training by partitioning models between clients and a server. However, under n

Beyond Accuracy: Decomposing the Reasoning Efficiency of LLMs

Model ReleasesDGX agent

arXiv:2602.09805v2 Announce Type: replace-cross Abstract: As reasoning LLMs increasingly trade tokens for accuracy through deliberation, search, and self-correction, a single accuracy score can no lon

Beyond Detection: A Structure-Aware Framework for Scene Text Tracking

Model ReleasesDGX agent

arXiv:2605.17270v1 Announce Type: new Abstract: Modern visual object trackers show impressive results on general targets, yet their performance drops substantially when dealing with scene text. Althou

Beyond Geometry: Efficient Topologically-Grounded Navigation in Complex 3D Environments

Model ReleasesDGX agent

arXiv:2605.17302v1 Announce Type: new Abstract: Ground robot navigation in complex 3D environments is often hindered by geometric ambiguity, where non-traversable structures such as furniture share lo

Beyond Neural Incompatibility: Cross-Scale Knowledge Transfer in Language Models through Latent Semantic Alignment

Model ReleasesDGX agent

arXiv:2510.24208v2 Announce Type: replace Abstract: Language Models (LMs) encode substantial knowledge in their parameters, yet it remains unclear how to transfer such knowledge in a fine-grained mann

Beyond Point-Wise Matching: Structural Representation Alignment for Accelerating Diffusion Transformers

Model ReleasesDGX agent

arXiv:2605.16949v1 Announce Type: new Abstract: Recent advances in Diffusion Transformers (DiTs) demonstrate that aligning noisy latent states with well-trained semantic features-as pioneered by Repre

Beyond Superficial Unlearning: Sharpness-Aware Robust Erasure of Hallucinations in Multimodal LLMs

Model ReleasesDGX agent

arXiv:2601.16527v2 Announce Type: replace-cross Abstract: Multimodal LLMs are powerful but prone to object hallucinations, which describe non-existent entities and harm reliability. While recent unlea

BioProAgent: Neuro-Symbolic Grounding for Constrained Scientific Planning

Model ReleasesDGX agent

arXiv:2603.00876v2 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated significant reasoning capabilities in scientific discovery but struggle to bridge the gap to physical

BlendedNet++: A dataset and benchmark for field-resolved aerodynamics and inverse design of blended wing body aircraft

Model ReleasesDGX agent

arXiv:2512.03280v2 Announce Type: replace-cross Abstract: The conceptual design of Blended Wing Body (BWB) aircraft is often constrained by the high computational cost of resolving complex aerodynamic

BoLT: A Benchmark to Democratize Black-box Optimization Research for Expensive LLM Tasks

Model ReleasesDGX agent

arXiv:2605.17000v1 Announce Type: cross Abstract: Optimization of LLM training and inference configurations, such as hyperparameters, data mixtures, and prompts, is critical to performance, but it is

Boundedly Rational Meta-Learning in Sequential Consumer Choice

Model ReleasesDGX agent

arXiv:2605.16532v1 Announce Type: new Abstract: Many consumer decisions are repeated choices under uncertainty. Standard models capture these decisions using Bayesian learning and dynamic programming:

Brain-inspired spike-timing plasticity for reliable label-efficient event-camera vision

Model ReleasesDGX agent

arXiv:2605.17686v1 Announce Type: new Abstract: Deploying event-camera object detectors is constrained by per-frame labeling requirements and GPU compute demands. This work introduces three local spik

Breaking Annotation Barriers: Generalized Video Quality Assessment via Ranking-based Self-Supervision

Model ReleasesDGX agent

arXiv:2505.03631v4 Announce Type: replace Abstract: Video quality assessment (VQA) is essential for quantifying perceptual quality in various video processing workflows, spanning from camera capture s

Bridging Data Trials and Task Barriers: A Unified Framework for Sketch Biometric Identification

Model ReleasesDGX agent

arXiv:2605.17367v1 Announce Type: new Abstract: Different from existing cross-modality identification tasks (e.g., heterogeneous face recognition, sketch re-identification, etc.), we introduce a novel

By now, you've probably heard about Gemini Omni, our new model designed to create anything from any input, starting with video. But... what'…

Model ReleasesDGX agent

Google AI announced Gemini Omni, a new multimodal model capable of generating diverse content types from various input formats, with initial focus on video generation capabilities. The model represent

CAM-Bench: A Benchmark for Computational and Applied Mathematics in Lean

Model ReleasesDGX agent

arXiv:2605.17255v1 Announce Type: new Abstract: Formal theorem-proving benchmarks enable mechanically verifiable evaluation of mathematical reasoning in large language models. However, existing benchm

Can Heterogeneous Language Models Be Fused?

Model ReleasesDGX agent

arXiv:2604.01674v2 Announce Type: replace Abstract: Model merging aims to integrate multiple expert models into a single model that inherits their complementary strengths without incurring the inferen

Can LLMs Generate and Solve Linguistic Olympiad Puzzles?

Model ReleasesDGX agent

arXiv:2509.21820v2 Announce Type: replace Abstract: In this paper, we introduce a combination of novel and exciting tasks: the solution and generation of linguistic puzzles. We focus on puzzles used i

Can LLMs Think Like Consumers? Benchmarking Crowd-Level Reaction Reconstruction with ConsumerSimBench

Model ReleasesDGX agent

arXiv:2605.17079v1 Announce Type: cross Abstract: LLMs are increasingly used as ``digital consumers'' to simulate public opinion, pre-test marketing decisions, and anticipate audience response. Howeve

Can These Views Be One Scene? Evaluating Multiview 3D Consistency when 3D Foundation Models Hallucinate

Model ReleasesDGX agent

arXiv:2605.18754v1 Announce Type: new Abstract: Multiview 3D evaluation assumes that the images being scored are observations of one static 3D scene. This assumption can fail in NVS and sparse-view re

CANSURF: An ASV-View Can Dataset and Benchmark for Detection and Tracking of Surface-Level Debris

Model ReleasesDGX agent

arXiv:2605.16774v1 Announce Type: cross Abstract: Surface-level marine debris remains a practical bottleneck for autonomous clean-up, where small, reflective targets (e.g., aluminum cans) must be dete

Can’t wait for Gemini Omni in @NotebookLM cinematic explainer videos 👀

Model ReleasesDGX agent

Emad Mostaque expressed anticipation for the integration of Google's Gemini Omni multimodal AI model into NotebookLM's cinematic explainer video generation features. The post suggests potential upcomi

CarbonScaling: Extending Neural Scaling Laws for Carbon Footprint in Large Language Models

Model ReleasesDGX agent

arXiv:2508.06524v2 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly follow neural scaling laws that tie performance gains to rapidly expanding computational budgets, ra

CAREBench: Evaluating LLMs' Emotion Understanding by Assessing Cognitive Appraisal Reasoning

Model ReleasesDGX agent

arXiv:2605.17176v1 Announce Type: new Abstract: Emotion understanding is a core capability for LLMs to interact effectively with humans, yet existing evaluation paradigms rely on discrete emotion labe

CasualSynth: Generating Structurally Sound Synthetic Data

Model ReleasesDGX agent

arXiv:2605.17528v1 Announce Type: cross Abstract: Large Language Models (LLMs) generate realistic synthetic data but offer no guarantee that their outputs respect the causal mechanisms governing the t

Causal Anomaly Detection for Lithium-Ion Battery Degradation

Model ReleasesDGX agent

arXiv:2605.17334v1 Announce Type: cross Abstract: Reliable early detection of lithium-ion battery degradation requires health indicators that are physically interpretable and computable from routine c

Causal Intervention-Based Memory Selection for Long-Horizon LLM Agents

Model ReleasesDGX agent

arXiv:2605.17641v1 Announce Type: new Abstract: Long-horizon LLM agents rely on persistent memory to support interactions across sessions, yet existing memory systems often retrieve context using sema

Causely: A Causal Intelligence Layer for Enterprise AI A Benchmark Study on SRE and Reliability Workflows

Model ReleasesDGX agent

arXiv:2605.18327v1 Announce Type: new Abstract: AI agents deployed into SRE workflows currently derive their understanding of environment state from raw observability telemetry at query time, paying a

CayleyPy RL: Pathfinding and Reinforcement Learning on Cayley Graphs

Model ReleasesDGX agent

arXiv:2502.18663v3 Announce Type: replace Abstract: This paper is the second in a series of studies on developing efficient artificial intelligence-based approaches to pathfinding on extremely large g

Cerebras is now running Kimi K2.6 – a trillion parameter model – in enterprise trials. At ~1,000 tokens/s, this is the fastest frontier mode…

Model ReleasesDGX agent

Cerebras is now running Kimi K2.6 – a trillion parameter model – in enterprise trials. At ~1,000 tokens/s, this is the fastest frontier model performance ever measured by Artificial Analysis @Artifici

CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?

Model ReleasesDGX agent

arXiv:2605.16679v1 Announce Type: cross Abstract: End-to-end automation of realistic healthcare operations stresses three capabilities underrepresented in current benchmarks: policy density, decisions

CLAP: Contrastive Latent-space Prompt Optimization for End-to-end Autonomous Driving

Model ReleasesDGX agent

arXiv:2605.17284v1 Announce Type: cross Abstract: End-to-end autonomous driving systems powered by Vision-Language-Action (VLA) models achieve strong performance on common driving scenarios, yet remai

ClawArena: Benchmarking AI Agents in Evolving Information Environments

Model ReleasesDGX agent

arXiv:2604.04202v2 Announce Type: replace-cross Abstract: AI agents deployed as persistent assistants must maintain correct beliefs as their information environment evolves. In practice, evidence is s

Closing the Gap at CRAC 2026: Two-Stage Adaptation for LLM-Based Multilingual Coreference Resolution

Model ReleasesDGX agent

arXiv:2605.16984v1 Announce Type: new Abstract: We present our submission to the LLM track of the 2026 Computational Models of Reference, Anaphora and Coreference (CRAC 2026) shared task. With an aver

Code-as-Room: Generating 3D Rooms from Top-Down View Images via Agentic Code Synthesis

Model ReleasesDGX agent

arXiv:2605.18451v1 Announce Type: new Abstract: Designing realistic and functional 3D indoor rooms is essential for a wide range of applications, including interior design, virtual reality, gaming, an

CommitDistill: A Lightweight Knowledge-Centric Memory Layer for Software Repositories

Model ReleasesDGX agent

arXiv:2605.18284v1 Announce Type: cross Abstract: Software repositories accumulate large amounts of unstructured knowledge in commit messages, pull-request discussions, and issue threads, but develope

CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection

Model ReleasesDGX agent

arXiv:2605.16839v1 Announce Type: new Abstract: Chunked prefill has become a widely adopted serving strategy for long-context large language models, but efficient attention computation in this regime

CompassAD: Intent-Driven 3D Affordance Grounding in Functionally Competing Objects

Model ReleasesDGX agent

arXiv:2604.02060v2 Announce Type: replace Abstract: When told to 'cut the cake,' a robot must choose the knife over nearby scissors, despite both objects affording the same cutting function. In real-w

Compounding Disadvantage: Auditing Intersectional Bias in LLM-Generated Explanations Across Indian and American STEM Education

Model ReleasesDGX agent

arXiv:2601.14506v3 Announce Type: replace-cross Abstract: Large language models are increasingly deployed in STEM education for personalized instruction and feedback across institutions in high- and l

Computer use turns Claude into an agent that can operate real UIs. New blog post on making it reliable in production: getting click accuracy…

Model ReleasesDGX agent

Computer use turns Claude into an agent that can operate real UIs. New blog post on making it reliable in production: getting click accuracy right, choosing thinking effort levels, keeping long sessio

Constrained Policy Optimization via Sampling-Based Weight-Space Projection

Model ReleasesDGX agent

arXiv:2512.13788v2 Announce Type: replace Abstract: Safety-critical learning requires policies that improve performance without leaving the safe operating regime. We study constrained policy learning

Context Memorization for Efficient Long Context Generation

Model ReleasesDGX agent

arXiv:2605.18226v1 Announce Type: cross Abstract: Modern large language model (LLM) applications increasingly rely on long conditioning prefixes to control model behavior at inference time. While pref

Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting

Model ReleasesDGX agent

arXiv:2508.04227v2 Announce Type: replace Abstract: Vision-language models (VLMs) and the recent surge of Multimodal Large Language Models (MLLMs) have revolutionized artificial intelligence with unpr

Continuous Diffusion Scales Competitively with Discrete Diffusion for Language

Model ReleasesDGX agent

arXiv:2605.18530v1 Announce Type: cross Abstract: While diffusion has drawn considerable recent attention from the language modeling community, continuous diffusion has appeared less scalable than dis

ContractBench: Can LLM Agents Preserve Observation Contracts?

Model ReleasesDGX agent

arXiv:2605.17281v1 Announce Type: cross Abstract: Tool-augmented LLM agents call APIs whose intermediate outputs, such as presigned URLs, session tokens, and OAuth state parameters, are observation co

ContraFix: Agentic Vulnerability Repair via Differential Runtime Evidence and Skill Reuse

Model ReleasesDGX agent

arXiv:2605.17450v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly used for automated vulnerability repair (AVR), where repository-level reasoning enables them to ins

Controlla: Learning Controllability via Graph-Constrained Latent Geometry

Model ReleasesDGX agent

arXiv:2605.16603v1 Announce Type: new Abstract: Controllable multimodal generation is commonly formulated as an inference-time conditioning problem using prompts, guidance, or auxiliary modules. While

Conv-FinRe: A Conversational and Longitudinal Benchmark for Utility-Grounded Financial Recommendation

Model ReleasesDGX agent

arXiv:2602.16990v2 Announce Type: replace Abstract: Most recommendation benchmarks evaluate how well a model imitates user behavior. In financial advisory, however, observed actions can be noisy or sh

Conversational editing: Gemini Omni allows you to edit your videos using natural language (like Nano Banana, but for video). So you can easi…

Model ReleasesDGX agent

Conversational editing: Gemini Omni allows you to edit your videos using natural language (like Nano Banana, but for video). So you can easily change your characters, settings, and styles by just desc

Coordinate Heterogeneity Governs Binary Quantization: From InfoNCE to Recall

Model ReleasesDGX agent

arXiv:2605.17524v1 Announce Type: new Abstract: Binary quantization (BQ) compresses high-dimensional embeddings into one or two bits per coordinate, enabling nearest neighbor search at extreme speed.

CooT: Learning to Coordinate In-Context with Coordination Transformers

Model ReleasesDGX agent

arXiv:2506.23549v3 Announce Type: replace Abstract: Effective coordination among unfamiliar partners remains a major challenge in multi-agent systems. Existing approaches, such as population-based met

Cost-aware Duration Prediction for Software Upgrades in Datacenters

Model ReleasesDGX agent

arXiv:2212.05155v2 Announce Type: replace-cross Abstract: Software upgrades are critical to maintaining server reliability in datacenters. While job duration prediction and scheduling have been extens

← Previous
1…233234235236237…377
Next →