AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,284 results
16 Apr 2026

What’s new with Google Data Cloud

Model ReleasesDGX agent

April 13 - April 17 We announced we are reintroducing Data Studio to play a significant role in the AI era, expanding from data visualizations and reports to host BigQuery conversational agents and da

When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning?

Model ReleasesDGX agent

arXiv:2503.23137v2 Announce Type: replace-cross Abstract: Understanding humor-particularly when it involves complex, contradictory narratives that require comparative reasoning-remains a significant c

Wordle 1,761 3/6 ⬛⬛⬛⬛🟨 ⬛⬛🟨🟨⬛ 🟩🟩🟩🟩🟩

Model ReleasesDGX agent

Anthropic's official X (Twitter) account shared a Wordle puzzle result for game number 1,761, solved in 3 out of 6 attempts. The post shows the characteristic colored tile pattern indicating the progr


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Wordle 1,762 4/6 ⬛🟨⬛⬛⬛ ⬛🟨⬛⬛⬛ 🟨⬛🟨🟨⬛ 🟩🟩🟩🟩🟩

Model ReleasesDGX agent

This post documents a Wordle game result where the player solved puzzle #1,762 in four attempts, using the standard color-coded feedback system (gray for incorrect letters, yellow for correct letters

Working Notes on Late Interaction Dynamics: Analyzing Targeted Behaviors of Late Interaction Models

Model ReleasesDGX agent

arXiv:2603.26259v2 Announce Type: replace-cross Abstract: While Late Interaction models exhibit strong retrieval performance, many of their underlying dynamics remain understudied, potentially hiding

WorkRB: A Community-Driven Evaluation Framework for AI in the Work Domain

Model ReleasesDGX agent

arXiv:2604.13055v1 Announce Type: new Abstract: Today's evolving labor markets rely increasingly on recommender systems for hiring, talent management, and workforce analytics, with natural language pr

You can also set your autocompact threshold yourself and effectively lower your context window if you'd prefer. For example, 400k context is…

Model ReleasesDGX agent

You can also set your autocompact threshold yourself and effectively lower your context window if you'd prefer. For example, 400k context is a good compromise: CLAUDE_CODE_AUTO_COMPACT_WINDOW=400000 c

15 Apr 2026

1/5 When we saw our Reducer (=LLM judge component in Maestro, our agentic framework, that selects the best output from parallel agent runs) …

Model ReleasesDGX agent

1/5 When we saw our Reducer (=LLM judge component in Maestro, our agentic framework, that selects the best output from parallel agent runs) consistently picking gold patches, we were sure Claude Opus

2/5 Turns out the model wasn't remembering the solution, but it was identifying 'gold-like' aesthetics like minimality & clarity. Total form…

Model ReleasesDGX agent

AI21 Labs shared findings indicating that their model does not simply memorize solutions but instead identifies and recognizes aesthetic qualities associated with high-quality outputs, such as minimal

A Benchmark for Evaluating Outcome-Driven Constraint Violations in Autonomous AI Agents

Model ReleasesDGX agent

arXiv:2512.20798v4 Announce Type: replace Abstract: As autonomous AI agents are deployed in high-stakes environments, ensuring their safety has become a paramount concern. Existing safety benchmarks p

A Foot Resistive Force Model for Legged Locomotion on Muddy Terrains

Model ReleasesDGX agent

arXiv:2604.12006v1 Announce Type: new Abstract: Legged robots face significant challenges in moving and navigating on deformable and highly yielding terrain such as mud. We present a resistive force m

A General Model for Deepfake Speech Detection: Diverse Bonafide Resources or Diverse AI-Based Generators

Model ReleasesDGX agent

arXiv:2603.27557v2 Announce Type: replace-cross Abstract: In this paper, we analyze two main factors of Bonafide Resource (BR) or AI-based Generator (AG) which affect the performance and the generalit

A Large-Scale Comparative Analysis of Imputation Methods for Single-Cell RNA Sequencing Data

Model ReleasesDGX agent

arXiv:2603.24626v2 Announce Type: replace-cross Abstract: Background: Single-cell RNA sequencing (scRNA-seq) enables gene expression profiling at cellular resolution but is inherently affected by spar

A Layer-wise Analysis of Supervised Fine-Tuning

Model ReleasesDGX agent

arXiv:2604.11838v1 Announce Type: cross Abstract: While critical for alignment, Supervised Fine-Tuning (SFT) incurs the risk of catastrophic forgetting, yet the layer-wise emergence of instruction-fol

A Sanity Check on Composed Image Retrieval

Model ReleasesDGX agent

arXiv:2604.12904v1 Announce Type: new Abstract: Composed Image Retrieval (CIR) aims to retrieve a target image based on a query composed of a reference image, and a relative caption that specifies the

ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning

Model ReleasesDGX agent

arXiv:2602.11236v2 Announce Type: replace-cross Abstract: Building general-purpose embodied agents across diverse hardware remains a central challenge in robotics, often framed as the ''one-brain, man

Adaptive Data Dropout: Towards Self-Regulated Learning in Deep Neural Networks

Model ReleasesDGX agent

arXiv:2604.12945v1 Announce Type: cross Abstract: Deep neural networks are typically trained by uniformly sampling large datasets across epochs, despite evidence that not all samples contribute equall

Adobe takes Creative Cloud into Claude Code-esque territory

Model ReleasesDGX agent

Adobe has introduced the **Firefly AI Assistant**, a new agentic tool that brings autonomous, multi-step workflow capabilities to Creative Cloud — drawing comparisons to AI coding agents like Claude C

AffectAgent: Collaborative Multi-Agent Reasoning for Retrieval-Augmented Multimodal Emotion Recognition

Model ReleasesDGX agent

arXiv:2604.12735v1 Announce Type: new Abstract: LLM-based multimodal emotion recognition relies on static parametric memory and often hallucinates when interpreting nuanced affective states. In this p

AISafetyBenchExplorer: A Metric-Aware Catalogue of AI Safety Benchmarks Reveals Fragmented Measurement and Weak Benchmark Governance

Model ReleasesDGX agent

arXiv:2604.12875v1 Announce Type: new Abstract: The rapid expansion of large language model (LLM) safety evaluation has produced a substantial benchmark ecosystem, but not a correspondingly coherent m

AlphaEval: Evaluating Agents in Production

Model ReleasesDGX agent

arXiv:2604.12162v1 Announce Type: new Abstract: The rapid deployment of AI agents in commercial settings has outpaced the development of evaluation methodologies that reflect production realities. Exi

Analyzing the Effect of Noise in LLM Fine-tuning

Model ReleasesDGX agent

arXiv:2604.12469v1 Announce Type: new Abstract: Fine-tuning is the dominant paradigm for adapting pretrained large language models (LLMs) to downstream NLP tasks. In practice, fine-tuning datasets may

Anthropic rolls out identity verification that may require Claude users to provide a government-issued photo ID and live selfie to access 'certain capabilities' (Jose Antonio Lanz/Decrypt)

Model ReleasesDGX agent

Jose Antonio Lanz / Decrypt: Anthropic rolls out identity verification that may require Claude users to provide a government-issued photo ID and live selfie to access “certain capabilities” — Anthropi

Anthropic’s Claude Code gets automated ‘routines’ and a desktop makeover

Model ReleasesDGX agent

Anthropic PBC is making it easier to automate tasks using Claude Code without relying on autonomous artificial intelligence agents with the launch of a new service called “routines.” The routines allo

AnyPoC: Universal Proof-of-Concept Test Generation for Scalable LLM-Based Bug Detection

Model ReleasesDGX agent

arXiv:2604.11950v1 Announce Type: cross Abstract: While recent LLM-based agents can identify many candidate bugs in source code, their reports remain static hypotheses that require manual validation,

ARC-AGI-3 has the lowest human bar of any AI benchmark out there. Almost all benchmarks require specialized knowledge that make them inacces…

Model ReleasesDGX agent

ARC-AGI-3 has the lowest human bar of any AI benchmark out there. Almost all benchmarks require specialized knowledge that make them inaccessible to 99%+ of humans (like, say SWE-Bench). ARC-AGI-3 is

Are Video Reasoning Models Ready to Go Outside?

Model ReleasesDGX agent

arXiv:2603.10652v2 Announce Type: replace-cross Abstract: In real-world deployment, vision-language models often encounter disturbances such as weather, occlusion, and camera motion. Under such condit

ARGOS: Who, Where, and When in Agentic Multi-Camera Person Search

Model ReleasesDGX agent

arXiv:2604.12762v1 Announce Type: cross Abstract: We introduce ARGOS, the first benchmark and framework that reformulates multi-camera person search as an interactive reasoning problem requiring an ag

ASTRA: Let Arbitrary Subjects Transform in Video Editing

Model ReleasesDGX agent

arXiv:2510.01186v2 Announce Type: replace Abstract: While existing video editing methods excel with single subjects, they struggle in dense, multi-subject scenes, frequently suffering from attention d

Banger paper from NVIDIA. Agentic reasoning needs models that are not just capable, but efficient at long-context inference. The agent model…

Model ReleasesDGX agent

Banger paper from NVIDIA. Agentic reasoning needs models that are not just capable, but efficient at long-context inference. The agent model layer is moving toward open, long-context, high-throughput

Been waiting a month for Anthropic to answer a simple usage question about Claude Code subscriptions Have I been ghosted

Model ReleasesDGX agent

Been waiting a month for Anthropic to answer a simple usage question about Claude Code subscriptions Have I been ghosted Can I get some questions answered by someone at Anthropic? 1. Can you use an OA

Benchmarking Deflection and Hallucination in Large Vision-Language Models

Model ReleasesDGX agent

arXiv:2604.12033v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) increasingly rely on retrieval to answer knowledge-intensive multimodal questions. Existing benchmarks overlook c

Benchmarking Foundation Models with Retrieval-Augmented Generation in Olympic-Level Physics Problem Solving

Model ReleasesDGX agent

arXiv:2510.00919v3 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) with foundation models has achieved strong performance across diverse tasks, but their capacity for exper

Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks

Model ReleasesDGX agent

arXiv:2604.12379v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly rely on explicit reasoning to solve coding tasks, yet evaluating the quality of this reasoning remains chall

Beyond Perception Errors: Semantic Fixation in Large Vision-Language Models

Model ReleasesDGX agent

arXiv:2604.12119v1 Announce Type: new Abstract: Large vision-language models (VLMs) often rely on familiar semantic priors, but existing evaluations do not cleanly separate perception failures from ru

Beyond Relevance: On the Relationship Between Retrieval and RAG Information Coverage

Model ReleasesDGX agent

arXiv:2603.08819v3 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) systems combine document retrieval with a generative model to address complex information seeking tasks l

Beyond Scores: Diagnostic LLM Evaluation via Fine-Grained Abilities

Model ReleasesDGX agent

arXiv:2604.12191v1 Announce Type: new Abstract: Current evaluations of large language models aggregate performance across diverse tasks into single scores. This obscures fine-grained ability variation

Beyond Single-Dimension Novelty: How Combinations of Theory, Method, and Results-based Novelty Shape Scientific Impact

Model ReleasesDGX agent

arXiv:2604.12471v1 Announce Type: cross Abstract: Scientific novelty drives advances at the research frontier, yet it is also associated with heightened uncertainty and potential resistance from incum

BID-LoRA: A Parameter-Efficient Framework for Continual Learning and Unlearning

Model ReleasesDGX agent

arXiv:2604.12686v1 Announce Type: cross Abstract: Recent advances in deep learning underscore the need for systems that can not only acquire new knowledge through Continual Learning (CL) but also remo

Bilevel Late Acceptance Hill Climbing for the Electric Capacitated Vehicle Routing Problem

Model ReleasesDGX agent

arXiv:2604.13013v1 Announce Type: new Abstract: This paper tackles the Electric Capacitated Vehicle Routing Problem (E-CVRP) through a bilevel optimization framework that handles routing and charging

Bipedal-Walking-Dynamics Model on Granular Terrains

Model ReleasesDGX agent

arXiv:2604.11981v1 Announce Type: new Abstract: Bipeds have demonstrated high agility and mobility in unstructured environments such as sand. The yielding of such granular media brings significant sin

btw the famous slack chart is slack propaganda and everyone who cites it is legally obligated to also link to @sophiebits

Model ReleasesDGX agent

btw the famous slack chart is slack propaganda and everyone who cites it is legally obligated to also link to @sophiebits Every time I see a tweet saying “I can vibe code this in a weekend” - I think

Budget blown on closed APIs is a solvable problem. Simply replace 20m of closed-source tokens with 1m on Minimax M2.7. Frontier performanc…

Model ReleasesDGX agent

Budget blown on closed APIs is a solvable problem. Simply replace 20m of closed-source tokens with 1m on Minimax M2.7. Frontier performance. 20x lower cost. No rate limits. If you're rethinking your A

Built GPT-2, Llama 3, and DeepSeek from scratch in PyTorch - open source code + book [p]

Model ReleasesDGX agent

A Reddit post on r/MachineLearning sharing an open-source project and accompanying book by Sebastian Raschka that walks through implementing GPT-2, Llama 3, and DeepSeek from scratch using PyTorch, wi

ByteDance launches its Seedance 2.0 video model to enterprise clients in 100+ countries, excluding the US amid legal disputes, after a February launch in China (Juro Osawa/The Information)

Model ReleasesDGX agent

Juro Osawa / The Information: ByteDance launches its Seedance 2.0 video model to enterprise clients in 100+ countries, excluding the US amid legal disputes, after a February launch in China — ByteDanc

Can AI Tools Transform Low-Demand Math Tasks? An Evaluation of Task Modification Capabilities

Model ReleasesDGX agent

arXiv:2604.12743v1 Announce Type: new Abstract: While recent research has explored AI tools' ability to classify the quality of mathematical tasks (arXiv:2603.03512), little is known about their capac

Can LLMs Beat Classical Hyperparameter Optimization Algorithms? A Study on autoresearch

Model ReleasesDGX agent

arXiv:2603.24647v4 Announce Type: replace Abstract: The autoresearch repository enables an LLM agent to optimize hyperparameters by editing training code directly. We use it as a testbed to compare cl

Capsule Security launches with $7M to secure AI agents at runtime

Model ReleasesDGX agent

Israeli agentic artificial intelligence security startup Capsule Security Ltd. today launched with 7 million in new funding to expand go-to-market efforts and accelerate product development across its

Chain-of-Models Pre-Training: Rethinking Training Acceleration of Vision Foundation Models

Model ReleasesDGX agent

arXiv:2604.12391v1 Announce Type: cross Abstract: In this paper, we present Chain-of-Models Pre-Training (CoM-PT), a novel performance-lossless training acceleration method for vision foundation model

Claude Opus 4.7 on Vertex AI

Model ReleasesDGX agent

Today, we’re announcing the general availability of Claude Opus 4.7 on Vertex AI. What’s new: Anthropic’s newest Opus model delivers advanced performance across coding, long-running agents, and profes

Climate Model Tuning with Online Synchronization-Based Parameter Estimation

Model ReleasesDGX agent

arXiv:2510.06180v2 Announce Type: replace-cross Abstract: In climate science, the tuning of climate models is a computationally intensive problem due to the combination of the high-dimensionality of t

Clustering-Enhanced Domain Adaptation for Cross-Domain Intrusion Detection in Industrial Control Systems

Model ReleasesDGX agent

arXiv:2604.12183v1 Announce Type: cross Abstract: Industrial control systems operate in dynamic environments where traffic distributions vary across scenarios, labeled samples are limited, and unknown

Clustering with Uniformity- and Neighbor-Based Random Geometric Graphs

Model ReleasesDGX agent

arXiv:2501.06268v3 Announce Type: replace Abstract: We propose a graph-based clustering method based on Cluster Catch Digraphs (CCDs) that extends their applicability to moderate-dimensional data sett

CoD-Lite: Real-Time Diffusion-Based Generative Image Compression

Model ReleasesDGX agent

arXiv:2604.12525v1 Announce Type: new Abstract: Recent advanced diffusion methods typically derive strong generative priors by scaling diffusion transformers. However, scaling fails to generalize when

CoDe-R: Refining Decompiler Output with LLMs via Rationale Guidance and Adaptive Inference

Model ReleasesDGX agent

arXiv:2604.12913v1 Announce Type: cross Abstract: Binary decompilation is a critical reverse engineering task aimed at reconstructing high-level source code from stripped executables. Although Large L

CodeSpecBench: Benchmarking LLMs for Executable Behavioral Specification Generation

Model ReleasesDGX agent

arXiv:2604.12268v1 Announce Type: cross Abstract: Large language models (LLMs) can generate code from natural language, but the extent to which they capture intended program behavior remains unclear.

CODESTRUCT: Code Agents over Structured Action Spaces

Model ReleasesDGX agent

arXiv:2604.05407v2 Announce Type: replace Abstract: LLM-based code agents treat repositories as unstructured text, applying edits through brittle string matching that frequently fails due to formattin

Comparison of low Steps, Klein 9b x Z image turbo x Ernie Turbo x Qwen 2512 8 Steps

Model ReleasesDGX agent

This r/StableDiffusion post presents a community-driven visual comparison of several modern, fast text-to-image diffusion models — Flux.2 Klein 9B (a smaller, faster distillation of Flux.2 Dev availab

Complementarity by Construction: A Lie-Group Approach to Solving Quadratic Programs with Linear Complementarity Constraints

Model ReleasesDGX agent

arXiv:2604.11991v1 Announce Type: new Abstract: Many problems in robotics require reasoning over a mix of continuous dynamics and discrete events, such as making and breaking contact in manipulation a

CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems

Model ReleasesDGX agent

arXiv:2604.12312v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly deployed as task-oriented agents in enterprise environments, ensuring their strict adherence to complex

← Previous
1…343344345346347…372
Next →