AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,136
  • Agents7,313
  • Applications5,230
  • Concepts5
  • Hardware1,765
  • Industry6,107
  • Local Ai4,758
  • Model Releases22,770
  • Research19,333
  • Safety12,890
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,136
  • Agents7,313
  • Applications5,230
  • Concepts5
  • Hardware1,765
  • Industry6,107
  • Local Ai4,758
  • Model Releases22,770
  • Research19,333
  • Safety12,890
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
85,136Total entries
1Added by human
85,135Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,778 results
Model Releases

BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law

DGX agent

arXiv:2605.28183v1 Announce Type: cross Abstract: We introduce the BenGER (Benchmark for German Law) dataset for evaluating LLM systems on subsumption-based legal reasoning in German law. The BenGER d

model-releasesarxiv-cs-ai
28 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Better Accuracies, Worse Reasoning: A Step-Level Audit of Medical Chain-of-Thought Distillation

DGX agent

arXiv:2605.28301v1 Announce Type: new Abstract: Chain-of-thought (CoT) distillation trains a smaller model to imitate a teacher's reasoning trace, but it is typically evaluated by final-answer metrics

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Beyond being fast, LiteParse is designed to provide highly accurate, semantically coherent text for LLM use. We benchmarked every open-sourc…

DGX agent

Beyond being fast, LiteParse is designed to provide highly accurate, semantically coherent text for LLM use. We benchmarked every open-source, model-free PDF parser on LLM QA tasks - from PyPDF to PyM

model-releasesjerry-liu--x
28 May 2026
Model Releases

Beyond Binary Moral Judgment: Modeling Ethical Pluralism in AI

DGX agent

arXiv:2605.28707v1 Announce Type: new Abstract: Critical decision-making in socially consequential spaces is increasingly involving AI systems at varying capacities. Yet, despite the ubiquity of auton

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring

DGX agent

arXiv:2502.05242v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are becoming increasingly capable, but the mechanisms of their thinking and decision-making processes remain uncl

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Beyond Model Ranking: Predictability-Aligned Evaluation for Time Series Forecasting

DGX agent

arXiv:2509.23074v3 Announce Type: replace-cross Abstract: In the era of increasingly complex AI models for time series forecasting, progress is often measured by marginal improvements on benchmark lea

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Beyond Motion Primitives: Behavioral Activity Recognition from Head-Mounted IMU

DGX agent

arXiv:2605.27464v1 Announce Type: cross Abstract: AR smart glasses need continuous behavioral context to offer proactive assistance, yet their most practical always-on sensor, the head-mounted Inertia

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Beyond One Path: Evaluating and Enhancing Divergent Thinking in Interactive LLM Agents

DGX agent

arXiv:2605.28465v1 Announce Type: new Abstract: Divergent thinking is a core dimension of creativity, yet existing evaluations of Large Language Models (LLMs) treat them as single-turn text generation

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

Bias Leaves a Gradient Trail: Label-Free Bias Identification via Gradient Probes on Concept Decompositions

DGX agent

arXiv:2605.28780v1 Announce Type: new Abstract: Vision classifiers can exploit spurious correlations, achieving high in-distribution accuracy yet failing under distribution shift. Existing approaches

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

Big migrations and refactors are some of a team's most important work, and the easiest to push off to a 'better time' since they'd tie up en…

DGX agent

Big migrations and refactors are some of a team's most important work, and the easiest to push off to a 'better time' since they'd tie up engineers for a quarter. With dynamic workflows, Claude can no

model-releasesboris-cherny--x
28 May 2026
Model Releases

Bilinear Coordinate Alignment for Training-Free Task-Vector Transfer

DGX agent

arXiv:2605.28444v1 Announce Type: new Abstract: Fine-tuning large-scale pre-trained models is a recent prevalent paradigm for adapting general representations to specialized tasks. However, when a new

model-releasesarxiv-cs-lg
28 May 2026
Model Releases

BioELX: Cross-lingual Biomedical Entity Linking via Alias-based Retrieval and LLM Ranking

DGX agent

arXiv:2605.27380v1 Announce Type: cross Abstract: Cross-lingual biomedical entity linking (BEL) maps mentions in any language to unique identifiers in a biomedical knowledge base (KB), supporting clin

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

BlazeEdit: Generalist Image Editing on Mobile Devices with Image-to-Image Diffusion Models

DGX agent

arXiv:2605.28067v1 Announce Type: new Abstract: The remarkable generation quality of modern diffusion models often comes at the cost of massive parameter counts, which necessitate server-side inferenc

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Bounded-Compute Multimodal Regression for Product-Rating Prediction

DGX agent

arXiv:2605.27737v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly attractive for multimodal quality assessment, but their default reliance on autoregressive text generatio

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

BREAKING: Anthropic just dropped Opus 4.8—and it is a MONSTER We've been testing for about a week @every and our verdict is they could've ju…

DGX agent

BREAKING: Anthropic just dropped Opus 4.8—and it is a MONSTER We've been testing for about a week @every and our verdict is they could've just called it Opus 5, it's that good. Here's our vibe check:

model-releasesthariq--x
28 May 2026
Model Releases

Bridging the Stability-Expressivity Gap: Synthetic Data Scaling and Preference Alignment for Low-Resource Spoken Language Models

DGX agent

arXiv:2605.27383v1 Announce Type: cross Abstract: Spoken Language Models (SLMs) have emerged as a promising paradigm for speech synthesis by bypassing explicit grapheme-to-phoneme pipelines. However,

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

BuddyBench: A Privacy-Constrained Multi-Task Benchmark for Pediatric Social-Communication Personalization

DGX agent

arXiv:2605.28089v1 Announce Type: new Abstract: BuddyBench introduces a privacy-constrained multi-task benchmark for pediatric social-communication personalization. Unlike existing neurodevelopmental

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Build a test suite that grows with your agent with dataset management in Amazon Bedrock AgentCore

DGX agent

Agent evaluation is most powerful when you combine fast-moving online signals with stable offline baselines. To understand whether your agent is truly improving over time, you need a fixed benchmark a

model-releasesaws-ml-blog
28 May 2026
Model Releases

Building Community-Centred NLP Resources for Puno Quechua

DGX agent

arXiv:2605.28253v1 Announce Type: new Abstract: The preservation of under-resourced languages requires digital tools and resources shaped by and for their speakers. We present the first dedicated ASR

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

Camellia: Benchmarking Cultural Biases in LLMs for Asian Languages

DGX agent

arXiv:2510.05291v2 Announce Type: replace Abstract: As Large Language Models (LLMs) develop stronger multilingual capabilities, their sensitivity to culturally diverse entities becomes increasingly im

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

Can Decision Trees Teach Large Language Models? Distilling Verbalized Knowledge for Molecular Property Prediction

DGX agent

arXiv:2603.12344v2 Announce Type: replace Abstract: Molecular Property Prediction (MPP) is a fundamental problem in drug discovery that has recently attracted growing attention. Large Language Models

model-releasesarxiv-cs-lg
28 May 2026
Model Releases

Can Large Language Models Handle Discourse Particles? A Case Study of Colloquial Malay

DGX agent

arXiv:2605.28782v1 Announce Type: new Abstract: Discourse particles, such as extit{well} and extit{kind of}, are crucial components that enable LLMs to ``speak'' more like humans. They are used to con

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

Can Segmentation Models Understand the World? Towards Proactive Affordance Reasoning via Visual Chain-of-Thought

DGX agent

arXiv:2605.27764v1 Announce Type: cross Abstract: Recent segmentation models couple large language models (LLMs) with mask decoders to ground complex language expressions into masks, yet their instruc

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

CAREF: Calibration-Aware Regularization for Explanation Faithfulness Without Rationale Supervision

DGX agent

arXiv:2605.27835v1 Announce Type: cross Abstract: We introduce CAREF, a parameter-efficient fine-tuning framework that jointly optimizes predictive accuracy and explanation faithfulness via calibratio

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

Category-Level 3D Correspondence in Camera Space via Morphable Object Priors

DGX agent

arXiv:2605.28257v1 Announce Type: new Abstract: Understanding 3D objects from images is fundamental to robotics and AR/VR applications. While recent work has made progress in category-level pose estim

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

CFDTwin: An open-source GUI and Python toolkit for POD-NN surrogate modeling of ANSYS Fluent simulations

DGX agent

arXiv:2605.27725v1 Announce Type: cross Abstract: High-fidelity computational fluid dynamics (CFD) is widely used for thermal-fluid design, but repeated CFD solves remain expensive for design optimiza

model-releasesarxiv-cs-lg
28 May 2026
Model Releases

ChildEval: When large language models meet children's personalities

DGX agent

arXiv:2605.27805v1 Announce Type: cross Abstract: While LLMs enable personalized chatbots, their effectiveness in child-centered personalization remains unclear, as systematic evaluation of child-spec

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Chinese Word Boundary Recovery through Character Alignment Projection

DGX agent

arXiv:2605.28128v1 Announce Type: new Abstract: Chinese word segmentation is especially fragile in non-standard text, where language learner errors and other character-level divergences disrupt the wo

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

Chirpy3D: Part-Aware Multi-View Diffusion for Creative Fine-Grained Object Generation

DGX agent

arXiv:2501.04144v3 Announce Type: replace Abstract: Understanding and generating the fine-grained structure of objects -- such as birds with species-specific beaks, wings, and tails -- is a long-stand

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

Chrome Enterprise rolls out AI agents and automation to streamline security management

DGX agent

Google LLC today launched new enhancements to Chrome Enterprise, the company’s enterprise version of its Chrome browser, designed to provide greater administrative and security control to information

model-releasessiliconangle
28 May 2026
Model Releases

CiteCheck: Retrieval-Grounded Detection of LLM Citation Hallucinations in Scientific Text

DGX agent

arXiv:2605.27700v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to generate scientific reports, but they can produce references that appear plausible while contain

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Claude Opus 4.8: 'a modest but tangible improvement'

DGX agent

Anthropic shipped Claude Opus 4.8 today. My favourite thing about it is this note in the release announcement: Users will find Opus 4.8 to be a modest but tangible improvement on its predecessor. Ther

model-releasessimon-willison
28 May 2026
Model Releases

Claude Opus 4.8 is now available for Max subscribers on Perplexity and Computer.

DGX agent

Claude Opus 4.8 has been made available to Max subscribers on the Perplexity platform and Computer application. This release expands access to Anthropic's Claude model through Perplexity's subscriptio

model-releasesperplexity--x
28 May 2026
Model Releases

Claude Opus 4.8 is now available in Cursor. On CursorBench, it's able to work much more efficiently than Opus 4.7. We've also found it to be…

DGX agent

Claude Opus 4.8 is now available as a model option in the Cursor code editor. According to Cursor's benchmarking, Opus 4.8 demonstrates improved efficiency compared to its predecessor Opus 4.7. The up

model-releasescursor--x
28 May 2026
Model Releases

Claude Opus 4.8 is now available in Windsurf and Devin CLI

DGX agent

Claude Opus 4.8 has been made available for use in both Windsurf (Cognition AI's code editor) and Devin CLI (their command-line interface). This release expands access to Anthropic's latest Claude mod

model-releasescognition-ai--x
28 May 2026
Model Releases

Claude Opus 4.8 is out today. It's our strongest coding model yet: up on SWE-bench Pro (from 64.3 to 69.2) and noticeably more honest about …

DGX agent

Claude Opus 4.8 is out today. It's our strongest coding model yet: up on SWE-bench Pro (from 64.3 to 69.2) and noticeably more honest about its own work. It tells you when it's unsure and catches its

model-releasesboris-cherny--x
28 May 2026
Model Releases

Claude’s new model is more ‘honest’ when it messes up

DGX agent

Anthropic is releasing Claude Opus 4.8 on Thursday, and the company is touting the model's 'honesty.' According to Anthropic, it trains 'all [its] models to be honest - for instance, to avoid making c

model-releasesthe-verge-ai
28 May 2026
Model Releases

ClinConsensus: A Physician-Calibrated Benchmark for Evaluating Clinical Rubric Coverage in Chinese Medical LLMs

DGX agent

arXiv:2603.02097v5 Announce Type: replace Abstract: Open-ended medical LLM evaluation remains weakly grounded in physician-calibrated coverage of clinically relevant response criteria, especially in l

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

ClinicalAgents: Multi-Agent Orchestration for Clinical Decision Making with Dual-Memory

DGX agent

arXiv:2603.26182v2 Announce Type: replace Abstract: While Large Language Models (LLMs) have demonstrated potential in healthcare, they often struggle with the complex, non-linear reasoning required fo

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

Code as a Weapon: A Consensus-Labeled Prompt Bank for Measuring Coding-Model Compliance with Malicious-Code Requests

DGX agent

arXiv:2605.28734v1 Announce Type: cross Abstract: A general-purpose language model that answers a harmful question returns text; a coding model that complies with a malicious request can return a work

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

CogPortrait: Fine-Grained Eye-Region Control in Portrait Animation via Hierarchical Agent Planning

DGX agent

arXiv:2605.28056v1 Announce Type: new Abstract: Portrait animation methods have achieved substantial visual quality and lip synchronization, but fine-grained manipulation of the eye region still faces

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification

DGX agent

arXiv:2508.21046v3 Announce Type: replace Abstract: Recent Vision-Language-Action (VLA) models built on pre-trained Vision-Language Models (VLMs) require extensive post-training, resulting in high com

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

Coherence Collapse: Diagnosing Why Code Agents Fail After Reaching the Right Code

DGX agent

arXiv:2603.24631v2 Announce Type: replace-cross Abstract: Code agents resolve 65-70% of SWE-bench Verified issues, but Pass@1 cannot tell us why the rest fail, and, as we show, capable-model failures

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Colosseum V2: Benchmarking Generalization for Vision Language Action Models

DGX agent

arXiv:2605.27759v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models demonstrate promising generalization in robotic manipulation, driven by advances in large-scale vision and language

model-releasesarxiv-cs-ro
28 May 2026
Model Releases

Comparative Analysis of Liquid Neural Networks and LSTM for Sequential Pattern Recognition: Robustness, Efficiency, and Clinical Utility

DGX agent

arXiv:2605.27467v1 Announce Type: cross Abstract: Traditional Recurrent Neural Networks (RNNs) and Long Short-Term Memory (LSTM) units operate on discrete time steps, often failing to capture the flui

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Confidence-Orchestrated Self-Evolution against Uncertain LLM Feedback

DGX agent

arXiv:2605.28010v1 Announce Type: new Abstract: Self-evolving large language models (LLMs) learn by generating their own training tasks and solutions, reducing reliance on human-curated supervision. H

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

ConRAG: Consensus-Driven Multi-View Retrieval for Multi-Hop Question Answering

DGX agent

arXiv:2605.28093v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) has emerged as a promising paradigm for enhancing large language models (LLMs) on multi-hop question answering (QA)

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

Conservative neural posterior estimation via distributionally robust training

DGX agent

arXiv:2605.28516v1 Announce Type: cross Abstract: Simulation-based inference with neural posterior estimation (NPE) often yields overconfident and unreliable posteriors under limited simulation budget

model-releasesarxiv-cs-lg
28 May 2026
← Previous
1…252253254255256…475
Next →