AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent
84,648Total entries
1Added by human
84,647Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,612 results
28 May 2026

Bilinear Coordinate Alignment for Training-Free Task-Vector Transfer

Model ReleasesDGX agent

arXiv:2605.28444v1 Announce Type: new Abstract: Fine-tuning large-scale pre-trained models is a recent prevalent paradigm for adapting general representations to specialized tasks. However, when a new

BioELX: Cross-lingual Biomedical Entity Linking via Alias-based Retrieval and LLM Ranking

Model ReleasesDGX agent

arXiv:2605.27380v1 Announce Type: cross Abstract: Cross-lingual biomedical entity linking (BEL) maps mentions in any language to unique identifiers in a biomedical knowledge base (KB), supporting clin

BlazeEdit: Generalist Image Editing on Mobile Devices with Image-to-Image Diffusion Models

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.28067v1 Announce Type: new Abstract: The remarkable generation quality of modern diffusion models often comes at the cost of massive parameter counts, which necessitate server-side inferenc

Bounded-Compute Multimodal Regression for Product-Rating Prediction

Model ReleasesDGX agent

arXiv:2605.27737v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly attractive for multimodal quality assessment, but their default reliance on autoregressive text generatio

BREAKING: Anthropic just dropped Opus 4.8—and it is a MONSTER We've been testing for about a week @every and our verdict is they could've ju…

Model ReleasesDGX agent

BREAKING: Anthropic just dropped Opus 4.8—and it is a MONSTER We've been testing for about a week @every and our verdict is they could've just called it Opus 5, it's that good. Here's our vibe check:

Bridging the Stability-Expressivity Gap: Synthetic Data Scaling and Preference Alignment for Low-Resource Spoken Language Models

Model ReleasesDGX agent

arXiv:2605.27383v1 Announce Type: cross Abstract: Spoken Language Models (SLMs) have emerged as a promising paradigm for speech synthesis by bypassing explicit grapheme-to-phoneme pipelines. However,

BuddyBench: A Privacy-Constrained Multi-Task Benchmark for Pediatric Social-Communication Personalization

Model ReleasesDGX agent

arXiv:2605.28089v1 Announce Type: new Abstract: BuddyBench introduces a privacy-constrained multi-task benchmark for pediatric social-communication personalization. Unlike existing neurodevelopmental

Build a test suite that grows with your agent with dataset management in Amazon Bedrock AgentCore

Model ReleasesDGX agent

Agent evaluation is most powerful when you combine fast-moving online signals with stable offline baselines. To understand whether your agent is truly improving over time, you need a fixed benchmark a

Building Community-Centred NLP Resources for Puno Quechua

Model ReleasesDGX agent

arXiv:2605.28253v1 Announce Type: new Abstract: The preservation of under-resourced languages requires digital tools and resources shaped by and for their speakers. We present the first dedicated ASR

Camellia: Benchmarking Cultural Biases in LLMs for Asian Languages

Model ReleasesDGX agent

arXiv:2510.05291v2 Announce Type: replace Abstract: As Large Language Models (LLMs) develop stronger multilingual capabilities, their sensitivity to culturally diverse entities becomes increasingly im

Can Decision Trees Teach Large Language Models? Distilling Verbalized Knowledge for Molecular Property Prediction

Model ReleasesDGX agent

arXiv:2603.12344v2 Announce Type: replace Abstract: Molecular Property Prediction (MPP) is a fundamental problem in drug discovery that has recently attracted growing attention. Large Language Models

Can Large Language Models Handle Discourse Particles? A Case Study of Colloquial Malay

Model ReleasesDGX agent

arXiv:2605.28782v1 Announce Type: new Abstract: Discourse particles, such as extit{well} and extit{kind of}, are crucial components that enable LLMs to ``speak'' more like humans. They are used to con

Can Segmentation Models Understand the World? Towards Proactive Affordance Reasoning via Visual Chain-of-Thought

Model ReleasesDGX agent

arXiv:2605.27764v1 Announce Type: cross Abstract: Recent segmentation models couple large language models (LLMs) with mask decoders to ground complex language expressions into masks, yet their instruc

CAREF: Calibration-Aware Regularization for Explanation Faithfulness Without Rationale Supervision

Model ReleasesDGX agent

arXiv:2605.27835v1 Announce Type: cross Abstract: We introduce CAREF, a parameter-efficient fine-tuning framework that jointly optimizes predictive accuracy and explanation faithfulness via calibratio

Category-Level 3D Correspondence in Camera Space via Morphable Object Priors

Model ReleasesDGX agent

arXiv:2605.28257v1 Announce Type: new Abstract: Understanding 3D objects from images is fundamental to robotics and AR/VR applications. While recent work has made progress in category-level pose estim

CFDTwin: An open-source GUI and Python toolkit for POD-NN surrogate modeling of ANSYS Fluent simulations

Model ReleasesDGX agent

arXiv:2605.27725v1 Announce Type: cross Abstract: High-fidelity computational fluid dynamics (CFD) is widely used for thermal-fluid design, but repeated CFD solves remain expensive for design optimiza

ChildEval: When large language models meet children's personalities

Model ReleasesDGX agent

arXiv:2605.27805v1 Announce Type: cross Abstract: While LLMs enable personalized chatbots, their effectiveness in child-centered personalization remains unclear, as systematic evaluation of child-spec

Chinese Word Boundary Recovery through Character Alignment Projection

Model ReleasesDGX agent

arXiv:2605.28128v1 Announce Type: new Abstract: Chinese word segmentation is especially fragile in non-standard text, where language learner errors and other character-level divergences disrupt the wo

Chirpy3D: Part-Aware Multi-View Diffusion for Creative Fine-Grained Object Generation

Model ReleasesDGX agent

arXiv:2501.04144v3 Announce Type: replace Abstract: Understanding and generating the fine-grained structure of objects -- such as birds with species-specific beaks, wings, and tails -- is a long-stand

Chrome Enterprise rolls out AI agents and automation to streamline security management

Model ReleasesDGX agent

Google LLC today launched new enhancements to Chrome Enterprise, the company’s enterprise version of its Chrome browser, designed to provide greater administrative and security control to information

CiteCheck: Retrieval-Grounded Detection of LLM Citation Hallucinations in Scientific Text

Model ReleasesDGX agent

arXiv:2605.27700v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to generate scientific reports, but they can produce references that appear plausible while contain

Claude Opus 4.8: 'a modest but tangible improvement'

Model ReleasesDGX agent

Anthropic shipped Claude Opus 4.8 today. My favourite thing about it is this note in the release announcement: Users will find Opus 4.8 to be a modest but tangible improvement on its predecessor. Ther

Claude Opus 4.8 is now available for Max subscribers on Perplexity and Computer.

Model ReleasesDGX agent

Claude Opus 4.8 has been made available to Max subscribers on the Perplexity platform and Computer application. This release expands access to Anthropic's Claude model through Perplexity's subscriptio

Claude Opus 4.8 is now available in Cursor. On CursorBench, it's able to work much more efficiently than Opus 4.7. We've also found it to be…

Model ReleasesDGX agent

Claude Opus 4.8 is now available as a model option in the Cursor code editor. According to Cursor's benchmarking, Opus 4.8 demonstrates improved efficiency compared to its predecessor Opus 4.7. The up

Claude Opus 4.8 is now available in Windsurf and Devin CLI

Model ReleasesDGX agent

Claude Opus 4.8 has been made available for use in both Windsurf (Cognition AI's code editor) and Devin CLI (their command-line interface). This release expands access to Anthropic's latest Claude mod

Claude Opus 4.8 is out today. It's our strongest coding model yet: up on SWE-bench Pro (from 64.3 to 69.2) and noticeably more honest about …

Model ReleasesDGX agent

Claude Opus 4.8 is out today. It's our strongest coding model yet: up on SWE-bench Pro (from 64.3 to 69.2) and noticeably more honest about its own work. It tells you when it's unsure and catches its

Claude’s new model is more ‘honest’ when it messes up

Model ReleasesDGX agent

Anthropic is releasing Claude Opus 4.8 on Thursday, and the company is touting the model's 'honesty.' According to Anthropic, it trains 'all [its] models to be honest - for instance, to avoid making c

ClinConsensus: A Physician-Calibrated Benchmark for Evaluating Clinical Rubric Coverage in Chinese Medical LLMs

Model ReleasesDGX agent

arXiv:2603.02097v5 Announce Type: replace Abstract: Open-ended medical LLM evaluation remains weakly grounded in physician-calibrated coverage of clinically relevant response criteria, especially in l

ClinicalAgents: Multi-Agent Orchestration for Clinical Decision Making with Dual-Memory

Model ReleasesDGX agent

arXiv:2603.26182v2 Announce Type: replace Abstract: While Large Language Models (LLMs) have demonstrated potential in healthcare, they often struggle with the complex, non-linear reasoning required fo

Code as a Weapon: A Consensus-Labeled Prompt Bank for Measuring Coding-Model Compliance with Malicious-Code Requests

Model ReleasesDGX agent

arXiv:2605.28734v1 Announce Type: cross Abstract: A general-purpose language model that answers a harmful question returns text; a coding model that complies with a malicious request can return a work

CogPortrait: Fine-Grained Eye-Region Control in Portrait Animation via Hierarchical Agent Planning

Model ReleasesDGX agent

arXiv:2605.28056v1 Announce Type: new Abstract: Portrait animation methods have achieved substantial visual quality and lip synchronization, but fine-grained manipulation of the eye region still faces

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification

Model ReleasesDGX agent

arXiv:2508.21046v3 Announce Type: replace Abstract: Recent Vision-Language-Action (VLA) models built on pre-trained Vision-Language Models (VLMs) require extensive post-training, resulting in high com

Coherence Collapse: Diagnosing Why Code Agents Fail After Reaching the Right Code

Model ReleasesDGX agent

arXiv:2603.24631v2 Announce Type: replace-cross Abstract: Code agents resolve 65-70% of SWE-bench Verified issues, but Pass@1 cannot tell us why the rest fail, and, as we show, capable-model failures

Colosseum V2: Benchmarking Generalization for Vision Language Action Models

Model ReleasesDGX agent

arXiv:2605.27759v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models demonstrate promising generalization in robotic manipulation, driven by advances in large-scale vision and language

Comparative Analysis of Liquid Neural Networks and LSTM for Sequential Pattern Recognition: Robustness, Efficiency, and Clinical Utility

Model ReleasesDGX agent

arXiv:2605.27467v1 Announce Type: cross Abstract: Traditional Recurrent Neural Networks (RNNs) and Long Short-Term Memory (LSTM) units operate on discrete time steps, often failing to capture the flui

Confidence-Orchestrated Self-Evolution against Uncertain LLM Feedback

Model ReleasesDGX agent

arXiv:2605.28010v1 Announce Type: new Abstract: Self-evolving large language models (LLMs) learn by generating their own training tasks and solutions, reducing reliance on human-curated supervision. H

ConRAG: Consensus-Driven Multi-View Retrieval for Multi-Hop Question Answering

Model ReleasesDGX agent

arXiv:2605.28093v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) has emerged as a promising paradigm for enhancing large language models (LLMs) on multi-hop question answering (QA)

Conservative neural posterior estimation via distributionally robust training

Model ReleasesDGX agent

arXiv:2605.28516v1 Announce Type: cross Abstract: Simulation-based inference with neural posterior estimation (NPE) often yields overconfident and unreliable posteriors under limited simulation budget

Continual Model Routing in Evolving Model Hubs

Model ReleasesDGX agent

arXiv:2605.28577v1 Announce Type: new Abstract: AI model hubs provide access to a rapidly growing collection of powerful pre-trained models, enabling off-the-shelf mixture-of-experts systems with diff

ConvMemory: A Lightweight Learned Memory Reranker, a Negative Attribution Result, and a Research-Preview Conflict Editor

Model ReleasesDGX agent

arXiv:2605.28062v1 Announce Type: new Abstract: We describe ConvMemory, a small 3.6M-parameter learned reranker for conversational long-term memory retrieval, trained with cross-encoder teacher superv

Cost-Sensitive Evaluation for Binary Classifiers

Model ReleasesDGX agent

arXiv:2510.22016v2 Announce Type: replace Abstract: Selecting an appropriate evaluation metric for classifiers is crucial for model comparison, parameter optimization, and deployment decisions, yet th

Cultural Binding Heads in Language Models

Model ReleasesDGX agent

arXiv:2605.28543v1 Announce Type: new Abstract: LLMs often default to equal treatment across cultural groups, even though context warrants differentiation: this is a lack of difference awareness. Usin

Cultural Fidelity in English-to-Hindi Translation: A Preservation-Fluency Frontier for Gender Recoverability

Model ReleasesDGX agent

arXiv:2605.27654v1 Announce Type: cross Abstract: Generative translation systems are cultural technologies because they decide how socially meaningful cues are rendered within culturally specific gram

CyberJurors: A Multi-Agent Simulation Task for E-Commerce Disputes Verdict

Model ReleasesDGX agent

arXiv:2605.28369v1 Announce Type: new Abstract: E-commerce platforms have begun recruiting crowdsourced jurors to adjudicate massive volumes of transaction disputes. Unlike formal legal judgment, E-co

Dark Quest II: A Wide-Coverage Neural Network Emulator of the Nonlinear Matter Power Spectrum Across Extended Cosmologies

Model ReleasesDGX agent

arXiv:2605.28596v1 Announce Type: cross Abstract: extsc{DarkEmulator2} is a neural network emulator of the nonlinear matter power spectrum in a nine-dimensional w_0 w_a nu o CDM parameter space, devel

Data-Efficient On-Policy Distillation for Automatic Speech Recognition

Model ReleasesDGX agent

arXiv:2605.28139v1 Announce Type: new Abstract: Building competitive automatic speech recognition (ASR) models usually requires large-scale au- dio supervision, which makes reproduction and specializa

Debate with Images: Detecting Deceptive Behaviors in Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2512.00349v2 Announce Type: replace Abstract: Are frontier AI systems becoming more capable? Certainly. Yet such progress is not an unalloyed blessing but rather a Trojan horse: behind their per

Decentralized Parameter-Free Online Learning with Compressed Gossip

Model ReleasesDGX agent

arXiv:2605.27831v1 Announce Type: new Abstract: We study decentralized online convex optimization when agents communicate over a graph and messages may be compressed. Classical decentralized online me

DecomposeRL: Learning to Ask Useful, Informative, and Diverse Questions for Semi-Supervised, Traceable Claim Verification

Model ReleasesDGX agent

arXiv:2605.27858v1 Announce Type: cross Abstract: Claim verification splits between end-to-end classifiers that are accurate but yields no inspectable traces, and decomposition-based methods produce i

Decoupled Training with Local Reinforcement Fine-Tuning in Federated Learning

Model ReleasesDGX agent

arXiv:2605.27900v1 Announce Type: new Abstract: Federated Learning (FL) with pre-trained Vision-Language Models (VLMs) has emerged as a promising paradigm for various downstream tasks. By leveraging i

Deepfake-Eval-2024: A Multi-Modal In-the-Wild Benchmark of Deepfakes Circulated in 2024

Model ReleasesDGX agent

arXiv:2503.02857v5 Announce Type: replace-cross Abstract: In the age of increasingly realistic generative AI, robust deepfake detection is essential for mitigating fraud and disinformation. While many

DeepSciVerify: Verifying Scientific Claim--Citation Alignment via LLM-Driven Evidence Escalation

Model ReleasesDGX agent

arXiv:2605.27710v1 Announce Type: new Abstract: Misalignment between claims and their cited evidence is a common failure mode in reports generated by large language models, limiting their reliability

Deformable Gaussian Occupancy: Decoupling Rigid and Nonrigid Motion with Factorized Distillation

Model ReleasesDGX agent

arXiv:2605.28587v1 Announce Type: new Abstract: Understanding dynamic 3D environments is essential for safe autonomous driving, particularly when reasoning about human-centric, nonrigid agents. Howeve

DEPART: DEcomposing PARiTy across Multilingual LLMs

Model ReleasesDGX agent

arXiv:2605.28163v1 Announce Type: cross Abstract: Multilingual Large Language Models (mLLMs) leaderboards report per-language accuracy but rarely explain why disparities emerge, leaving systemic biase

Detection Without Correction: A Two-Parameter Decomposition of Multi-Stage LLM Pipelines

Model ReleasesDGX agent

arXiv:2605.27559v1 Announce Type: cross Abstract: Multi-stage LLM pipelines that perform multi-agent debate, intrinsic self-correction, or retrieval-augmented verification exhibit puzzling aggregate b

'Developers can update Claude’s instructions mid-task without breaking the prompt cache or routing the update through a user turn' wtf? how?…

Model ReleasesDGX agent

'Developers can update Claude’s instructions mid-task without breaking the prompt cache or routing the update through a user turn' wtf? how?? Introducing Claude Opus 4.8: it builds on Opus 4.7 with sh

Did this actually happen? It seems very suspicious.

Model ReleasesDGX agent

Did this actually happen? It seems very suspicious. 'An AI consultant tells Axios one of their clients recently spent half a billion dollars in a single month after failing to put usage limits on Clau

Differential syntactic and semantic encoding in LLMs

Model ReleasesDGX agent

arXiv:2601.04765v4 Announce Type: replace-cross Abstract: We study how syntactic and semantic information is encoded in inner layer representations of Large Language Models (LLMs), focusing on the ver

Diffusion-Based Ukrainian Handwritten Text Generation with Cross-Domain Style Transfer

Model ReleasesDGX agent

arXiv:2605.27487v1 Announce Type: cross Abstract: Handwritten text generation (HTG) conditioned on writer style has been widely studied for Latin scripts, but remains underexplored for low-resource an

Dimensionality Reduction for Robust Federated Learning: A Theoretical Analysis and Convergence Guarantee

Model ReleasesDGX agent

arXiv:2605.28335v1 Announce Type: new Abstract: Federated Learning (FL) enables multiple clients to collaboratively train models without sharing raw data, but it is highly vulnerable to Byzantine atta

← Previous
1…198199200201202…377
Next →