AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent
84,648Total entries
1Added by human
84,647Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,612 results
28 May 2026

A Fresh Look at Lamarckian Evolution and the Baldwin Effect

Model ReleasesDGX agent

arXiv:2605.28703v1 Announce Type: cross Abstract: Baldwinian and Lamarckian evolution have existed for a long time in evolutionary algorithms (EAs) without ever dominating the academic literature or p

A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks

Model ReleasesDGX agent

arXiv:2605.28556v1 Announce Type: new Abstract: As agent capabilities advance, existing benchmarks, such as au^2-Bench, are becoming increasingly saturated. Yet constructing new benchmark tasks remain

A Multi-dimensional Framework for Evaluating Generalization in EEG Foundation Models

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.28563v1 Announce Type: cross Abstract: Evaluating foundation models under appropriate adaptation settings is essential for understanding the quality and transferability of the learned repre

A Query Engine for the Agents

Model ReleasesDGX agent

arXiv:2605.27785v1 Announce Type: new Abstract: The fastest-growing data in production today is unstructured text: agent traces, chat logs, reasoning chains, model outputs. People want to analyze it,

A Simple State Space Model Excels at Multivariate Time Series Classification

Model ReleasesDGX agent

arXiv:2605.27406v1 Announce Type: new Abstract: Structured state space models (SSMs) have recently emerged as a promising foundation for sequence modeling, with Mamba-based architectures demonstrating

A Unified Framework for the Evaluation of LLM Agentic Capabilities

Model ReleasesDGX agent

arXiv:2605.27898v1 Announce Type: new Abstract: As LLMs are increasingly deployed as agents, reliable assessment of their agentic capabilities has become essential. However, reported benchmark scores

AdaDPO: Self-Adaptive Direct Preference Optimization with Balanced Gradient Updates

Model ReleasesDGX agent

arXiv:2605.28440v1 Announce Type: new Abstract: DPO has become a widely adopted alternative to RLHF for aligning LLMs with human preferences, eliminating the need for a separate reward model or RL loo

Adaptive Bandit Algorithms for Contextual Matching Markets

Model ReleasesDGX agent

arXiv:2605.28290v1 Announce Type: new Abstract: We study bandit learning in matching markets, where players and arms constitute the two market sides, and the players' utilities are linear in the arm c

Adaptive Cost-Efficient Evaluation for Reliable Patent Claim Generation

Model ReleasesDGX agent

arXiv:2604.04295v3 Announce Type: replace Abstract: Automated patent claim validation demands low error tolerance. However, existing approaches face a rigidity-resource dilemma: lightweight encoders c

Adaptive Reservoir Computing for Multi-Scenario Chaotic System Forecasting

Model ReleasesDGX agent

arXiv:2605.28145v1 Announce Type: new Abstract: We present an adaptive reservoir computing framework for the CTF-4-Science Lorenz benchmark, which evaluates machine learning models across twelve disti

Adversarial Fine-tuning of Compressed Neural Networks for Joint Improvement of Robustness and Efficiency

Model ReleasesDGX agent

arXiv:2403.09441v2 Announce Type: replace Abstract: As deep learning (DL) models are increasingly being integrated into our everyday lives, ensuring their safety by making them robust against adversar

AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens

Model ReleasesDGX agent

arXiv:2512.17375v2 Announce Type: replace-cross Abstract: LLM-as-a-Judge systems supply the reward signal in modern RLHF and RLVR pipelines, but their binary verdict reduces to a single linear readout

Agentic Active Omni-Modal Perception for Multi-Hop Audio-Visual Reasoning

Model ReleasesDGX agent

arXiv:2605.28192v1 Announce Type: new Abstract: Multi-hop audio-visual reasoning remains challenging for Omni-LLMs, as relevant evidence is often sparse, temporally dispersed, and distributed across b

Agentic Separation Logic Specification Synthesis

Model ReleasesDGX agent

arXiv:2605.27531v1 Announce Type: cross Abstract: Specification synthesis, the task of automatically inferring formal specifications from program implementations and natural language, is important for

AI in SRE: Where and how Google is deploying agentic AI to improve operations

Model ReleasesDGX agent

Since its inception over 20 years ago, Google has used Site Reliability Engineering (SRE) to keep services like Search, Gmail, Maps, YouTube and Google Cloud reliable and highly available, adhering to

AI researchers ran 15-day simulations of worlds governed by different AI models: Claude Sonnet 4.6 recorded no crimes, while Gemini 3 Flash had the most at 683 (Jake Angelo/Fortune)

Model ReleasesDGX agent

Jake Angelo / Fortune: AI researchers ran 15-day simulations of worlds governed by different AI models: Claude Sonnet 4.6 recorded no crimes, while Gemini 3 Flash had the most at 683 — Imagine a world

aight bro nvm the bouncer is an opp just show up whenever lol

Model ReleasesDGX agent

This appears to be a casual, informal social media post using slang terminology, likely discussing plans to attend an event or venue while making light of potential conflicts with a bouncer. The post

Aligning Language Model Benchmarks with Pairwise Preferences

Model ReleasesDGX agent

arXiv:2602.02898v2 Announce Type: replace Abstract: Language model benchmarks are pervasive and computationally-efficient proxies for real-world performance. However, many recent works find that bench

AlphaForgeBench: Benchmarking End-to-End Trading Strategy Design with Large Language Models

Model ReleasesDGX agent

arXiv:2602.18481v2 Announce Type: replace-cross Abstract: The rapid advancement of Large Language Models (LLMs) has led to a surge of financial benchmarks, evolving from static knowledge evaluation to

AlphaTransit: Learning to Design City-scale Transit Routes

Model ReleasesDGX agent

arXiv:2605.28730v1 Announce Type: new Abstract: Designing a transit network requires many sequential route extension decisions, but their quality is often visible only after the full network is assemb

Also out today: You can now directly configure the effort level and adaptive thinking in Code (/effort) and Cowork! Effort allows you to tun…

Model ReleasesDGX agent

Also out today: You can now directly configure the effort level and adaptive thinking in Code (/effort) and Cowork! Effort allows you to tune Claude's intelligence vs token spend, trading off capabili

An Enhanced Large Neighborhood Search Approach for the Capacitated Facility Location Problem with Incompatible Customers

Model ReleasesDGX agent

arXiv:2605.28337v1 Announce Type: new Abstract: A new variant of the classic capacitated facility location problem, which considers incompatibilities between customers, has recently been introduced in

Analyzing Quality-Latency-Resource Trade-offs in a Technical Documentation RAG Assistant Using LoRA Adaptation

Model ReleasesDGX agent

arXiv:2605.28222v1 Announce Type: new Abstract: We study quality-latency-resource trade-offs in a documentation-grounded retrieval-augmented generation (RAG) system that uses Low-Rank Adaptation (LoRA

AndroidDaily: A Verifiable Benchmark for Mobile GUI Agents on Real-World Closed-Source Applications

Model ReleasesDGX agent

arXiv:2605.27761v1 Announce Type: new Abstract: The rapid development of GUI foundation models and mobile GUI agents has spurred numerous evaluation benchmarks, yet most rely on simulated environments

Announcing the newest cohort of the Google for Startups Accelerator: Middle East, North Africa & Turkey

Model ReleasesDGX agent

Google’s mission is to organize the world’s information and make it universally accessible. In high-growth, technically ambitious markets like the Middle East, North Africa, and Türkiye (MENA-T), we f

Anthropic adds dynamic workflows to Claude Code, enabling hundreds of subagents to run in parallel for complex engineering tasks such as framework migrations (Claude)

Model ReleasesDGX agent

Claude: Anthropic adds dynamic workflows to Claude Code, enabling hundreds of subagents to run in parallel for complex engineering tasks such as framework migrations — Early access users and teams ins

Anthropic says it expects Mythos-class models to be available to all customers 'in the coming weeks' following the development of stronger safeguards (Madison Mills/Axios)

Model ReleasesDGX agent

Madison Mills / Axios: Anthropic says it expects Mythos-class models to be available to all customers “in the coming weeks” following the development of stronger safeguards — Anthropic released Claude

Apple Intelligence Foundation Language Models

Model ReleasesDGX agent

arXiv:2407.21075v2 Announce Type: replace Abstract: We present foundation language models developed to power Apple Intelligence features, including a ~3 billion parameter model designed to run efficie

Apple working to cram massive Gemini model into iPhone to power new Siri

Model ReleasesDGX agent

Apple is reportedly working to distill knowledge and skills from Google's larger Gemini model into a smaller version that could run on iPhones. The new Siri will use a tiered system where simple tasks

Are Large Pre-trained Vision Language Models Effective Construction Safety Inspectors?

Model ReleasesDGX agent

arXiv:2508.11011v2 Announce Type: replace Abstract: Construction safety inspections typically involve a human inspector identifying safety concerns on-site. With the rise of powerful Vision Language M

Argument Quality Assessment with Large Language Models: A Pairwise Bradley-Terry Approach

Model ReleasesDGX agent

arXiv:2605.28313v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities in tasks related to reasoning and judgment. However, assessing the quality of arg

Ariel-ML: Computing Parallelization with Embedded Rust for Neural Networks on Heterogeneous Multi-core Microcontrollers

Model ReleasesDGX agent

arXiv:2512.09800v2 Announce Type: replace Abstract: Low-power microcontroller (MCU) hardware is currently evolving from single-core architectures to predominantly multi-core architectures. In parallel

As Anthropic launches Claude Opus 4.8, it raises $65B in new funding

Model ReleasesDGX agent

Anthropic PBC today introduced a new large language model, Claude Opus 4.8, that’s significantly better than its predecessor at complex coding tasks. The company announced the LLM alongside another ma

Ask Now, Use Later: Benchmarking the Proactivity Gap in Long-Lived LLM Agents

Model ReleasesDGX agent

arXiv:2605.28108v1 Announce Type: new Abstract: A long-lived LLM agent, such as OpenClaw, earns its value by acting on a user's preferences and constraints across sessions, not just the current reques

AssertLLM2: A Comprehensive LLM Benchmark for Assertion Generation from Design Specifications

Model ReleasesDGX agent

arXiv:2605.27472v1 Announce Type: cross Abstract: Assertion-based verification (ABV) is a cornerstone of modern hardware design, yet manually translating design intent into formal SystemVerilog Assert

Assessing Factual Music Comprehension in Large Audio Language Models

Model ReleasesDGX agent

arXiv:2511.05550v2 Announce Type: replace-cross Abstract: Large audio language models (LALMs) leverage multimodal representations to generate open-ended answers to natural language queries about audio

ASTRA: Communication-Efficient Acceleration for Multi-Device Transformer Inference

Model ReleasesDGX agent

arXiv:2505.19342v2 Announce Type: replace-cross Abstract: Multi-device inference can reduce Transformer latency by parallelizing computation. However, existing methods require high inter-device bandwi

Asynchronous Remote Sensing Time-Series Fusion for Cloud Removal and Anytime Reconstruction

Model ReleasesDGX agent

arXiv:2605.27726v1 Announce Type: new Abstract: Frequent cloud cover severely limits the usability of Sentinel-2 (S2) optical time series for Earth surface monitoring. Sentinel-1 (S1) SAR provides all

AsyncTool: Evaluating the Asynchronous Function Calling Capability under Multi-Task Scenarios

Model ReleasesDGX agent

arXiv:2605.27995v1 Announce Type: new Abstract: Large language model (LLM)-based agents have shown strong capabilities in using external tools to solve complex tasks. However, existing evaluations oft

ATLAS: All-round Testing of Long-context Abilities across Scales

Model ReleasesDGX agent

arXiv:2605.28079v1 Announce Type: new Abstract: Long-context language models now advertise context windows up to millions of tokens, yet evaluations typically report a single length or a narrow task f

Augmenting Attention with Exponentially Decaying Memory Improves Query-Aware KV Sparsity

Model ReleasesDGX agent

arXiv:2605.28640v1 Announce Type: new Abstract: Efficient inference is critical for long-context language models, where attention computation and KV-cache access dominate the cost. Recent work RAT+, i

Automating Formal Verification with Agent-Guided Tree Search

Model ReleasesDGX agent

arXiv:2605.27485v1 Announce Type: cross Abstract: Formal verification offers a path to provably correct software, but writing verified code remains expensive enough that the technique is rarely used i

Bandwidth-Efficient and Privacy-Preserving Edge-Cloud Many-to-Many Speech Translation

Model ReleasesDGX agent

arXiv:2605.28642v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have demonstrated significant potential for speech-to-text translation (S2TT). However, existing deployment par

Bayesian Optimization Parameter Tuning Framework for a Lyapunov Based Path Following Controller

Model ReleasesDGX agent

arXiv:2512.12649v2 Announce Type: replace Abstract: Parameter tuning in real-world experiments is constrained by the limited evaluation budget available on hardware. The path-following controller stud

Benchmarking AI for low-resource contexts: Thinking beyond leaderboards

Model ReleasesDGX agent

arXiv:2605.28508v1 Announce Type: new Abstract: Existing AI evaluation practices often fail to capture how systems actually perform in low-resource environments, where operational constraints shape us

Benchmarking and Mechanistic Analysis of Vision-Language Models for Cross-Depiction Assembly Instruction Alignment

Model ReleasesDGX agent

arXiv:2604.00913v2 Announce Type: replace-cross Abstract: 2D assembly diagrams are often abstract and hard to follow, creating a need for intelligent assistants that can monitor progress, detect error

Benchmarking Fairness in Spiking Neural Networks: Data Bias, Spurious Features, and Hardware Effects

Model ReleasesDGX agent

arXiv:2605.27407v1 Announce Type: cross Abstract: Evaluating fairness in Spiking Neural Networks (SNNs) demands rigorous benchmarks that reflect real-world complexities, yet existing assessments remai

Benchmarking Inductive Biases for Multivariate Time-Series Anomaly Detection with a Robust Multi-View Channel-Graph Detector

Model ReleasesDGX agent

arXiv:2605.28103v1 Announce Type: new Abstract: We present a unified experiment, analysis, and benchmark study of multivariate time-series (MTS) anomaly detection. Ten family-representative detectors

Benchmarking Ultrasound Foundation Models for Fetal Plane Classification

Model ReleasesDGX agent

arXiv:2605.27796v1 Announce Type: cross Abstract: Ultrasound is widely used in obstetric care due to its safety, accessibility, and real-time imaging. However, interpretation remains operator-dependen

Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systems

Model ReleasesDGX agent

arXiv:2605.27492v1 Announce Type: cross Abstract: LLM agents are rapidly evolving from coding assistants into autonomous software engineering systems. However, existing evaluation methodologies remain

BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law

Model ReleasesDGX agent

arXiv:2605.28183v1 Announce Type: cross Abstract: We introduce the BenGER (Benchmark for German Law) dataset for evaluating LLM systems on subsumption-based legal reasoning in German law. The BenGER d

Better Accuracies, Worse Reasoning: A Step-Level Audit of Medical Chain-of-Thought Distillation

Model ReleasesDGX agent

arXiv:2605.28301v1 Announce Type: new Abstract: Chain-of-thought (CoT) distillation trains a smaller model to imitate a teacher's reasoning trace, but it is typically evaluated by final-answer metrics

Beyond being fast, LiteParse is designed to provide highly accurate, semantically coherent text for LLM use. We benchmarked every open-sourc…

Model ReleasesDGX agent

Beyond being fast, LiteParse is designed to provide highly accurate, semantically coherent text for LLM use. We benchmarked every open-source, model-free PDF parser on LLM QA tasks - from PyPDF to PyM

Beyond Binary Moral Judgment: Modeling Ethical Pluralism in AI

Model ReleasesDGX agent

arXiv:2605.28707v1 Announce Type: new Abstract: Critical decision-making in socially consequential spaces is increasingly involving AI systems at varying capacities. Yet, despite the ubiquity of auton

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring

Model ReleasesDGX agent

arXiv:2502.05242v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are becoming increasingly capable, but the mechanisms of their thinking and decision-making processes remain uncl

Beyond Model Ranking: Predictability-Aligned Evaluation for Time Series Forecasting

Model ReleasesDGX agent

arXiv:2509.23074v3 Announce Type: replace-cross Abstract: In the era of increasingly complex AI models for time series forecasting, progress is often measured by marginal improvements on benchmark lea

Beyond Motion Primitives: Behavioral Activity Recognition from Head-Mounted IMU

Model ReleasesDGX agent

arXiv:2605.27464v1 Announce Type: cross Abstract: AR smart glasses need continuous behavioral context to offer proactive assistance, yet their most practical always-on sensor, the head-mounted Inertia

Beyond One Path: Evaluating and Enhancing Divergent Thinking in Interactive LLM Agents

Model ReleasesDGX agent

arXiv:2605.28465v1 Announce Type: new Abstract: Divergent thinking is a core dimension of creativity, yet existing evaluations of Large Language Models (LLMs) treat them as single-turn text generation

Bias Leaves a Gradient Trail: Label-Free Bias Identification via Gradient Probes on Concept Decompositions

Model ReleasesDGX agent

arXiv:2605.28780v1 Announce Type: new Abstract: Vision classifiers can exploit spurious correlations, achieving high in-distribution accuracy yet failing under distribution shift. Existing approaches

Big migrations and refactors are some of a team's most important work, and the easiest to push off to a 'better time' since they'd tie up en…

Model ReleasesDGX agent

Big migrations and refactors are some of a team's most important work, and the easiest to push off to a 'better time' since they'd tie up engineers for a quarter. With dynamic workflows, Claude can no

← Previous
1…197198199200201…377
Next →