AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,106 results
10 Apr 2026

Anthropic and OpenAI target big businesses with enterprise-grade controls and lower pricing

Model ReleasesDGX agent

Artificial intelligence leaders Anthropic PBC and OpenAI Group PBC are stepping up their efforts to compete for the enterprise, making their most advanced agentic tools more accessible to the largest

Anthropic tries to keep its new AI model away from cyberattackers as enterprises look to tame AI chaos

Model ReleasesDGX agent

Sure, at some point quantum computing may break data encryption — but well before that, artificial intelligence models already seem likely to wreak havoc. That became starkly apparent this week when A

Anthropic's Claude Mythos isn't a sentient super-hacker, it's a sales pitch — claims of 'thousands' of severe zero-days rely on just 198 man…


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases
DGX agent

Anthropic's Claude Mythos isn't a sentient super-hacker, it's a sales pitch — claims of 'thousands' of severe zero-days rely on just 198 manual reviews https://www.tomshardware.com/tech-industry/artif

Anticipating tipping in spatiotemporal systems with machine learning

Model ReleasesDGX agent

arXiv:2604.06454v1 Announce Type: cross Abstract: In nonlinear dynamical systems, tipping refers to a critical transition from one steady state to another, typically catastrophic, steady state, often

Apple: Toward General Active Perception via Reinforcement Learning

Model ReleasesDGX agent

arXiv:2505.06182v5 Announce Type: replace-cross Abstract: Active perception is a fundamental skill that enables us humans to deal with uncertainty in our inherently partially observable environment. F

Are Stochastic Multi-objective Bandits Harder than Single-objective Bandits?

Model ReleasesDGX agent

arXiv:2604.07096v1 Announce Type: new Abstract: Multi-objective bandits have attracted increasing attention because of their broad applicability and mathematical elegance, where the reward of each arm

arXiv2Table: Toward Realistic Benchmarking and Evaluation for LLM-Based Literature-Review Table Generation

Model ReleasesDGX agent

arXiv:2504.10284v5 Announce Type: replace Abstract: Literature review tables are essential for summarizing and comparing collections of scientific papers. In this paper, we study the automatic generat

ASBench: Image Anomalies Synthesis Benchmark for Anomaly Detection

Model ReleasesDGX agent

arXiv:2510.07927v2 Announce Type: replace Abstract: Anomaly detection plays a pivotal role in manufacturing quality control, yet its application is constrained by limited abnormal samples and high man

Asking like Socrates: Socrates helps VLMs understand remote sensing images

Model ReleasesDGX agent

arXiv:2511.22396v2 Announce Type: replace-cross Abstract: Recent multimodal reasoning models, inspired by DeepSeek-R1, have significantly advanced vision-language systems. However, in remote sensing (

Asymptotic-Preserving Neural Networks for Viscoelastic Parameter Identification in Multiscale Blood Flow Modeling

Model ReleasesDGX agent

arXiv:2604.06287v1 Announce Type: new Abstract: Mathematical models and numerical simulations offer a non-invasive way to explore cardiovascular phenomena, providing access to quantities that cannot b

ATANT: An Evaluation Framework for AI Continuity

Model ReleasesDGX agent

arXiv:2604.06710v1 Announce Type: new Abstract: We present ATANT (Automated Test for Acceptance of Narrative Truth), an open evaluation framework for measuring continuity in AI systems: the ability to

ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis

Model ReleasesDGX agent

arXiv:2604.02022v2 Announce Type: replace Abstract: Evaluating the safety of LLM-based agents is increasingly important because risks in realistic deployments often emerge over multi-step interactions

AtlasOCR: Building the First Open-Source Darija OCR Model with Vision Language Models

Model ReleasesDGX agent

arXiv:2604.08070v1 Announce Type: new Abstract: Darija, the Moroccan Arabic dialect, is rich in visual content yet lacks specialized Optical Character Recognition (OCR) tools. This paper introduces At

Automating Database-Native Function Code Synthesis with LLMs

Model ReleasesDGX agent

arXiv:2604.06231v1 Announce Type: cross Abstract: Database systems incorporate an ever-growing number of functions in their kernels (a.k.a., database native functions) for scenarios like new applicati

AutoReproduce: Automatic AI Experiment Reproduction with Paper Lineage

Model ReleasesDGX agent

arXiv:2505.20662v3 Announce Type: replace Abstract: Efficient reproduction of research papers is pivotal to accelerating scientific progress. However, the increasing complexity of proposed methods oft

AV-SQL: Decomposing Complex Text-to-SQL Queries with Agentic Views

Model ReleasesDGX agent

arXiv:2604.07041v1 Announce Type: cross Abstract: Text-to-SQL is the task of translating natural language queries into executable SQL for a given database, enabling non-expert users to access structur

AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation

Model ReleasesDGX agent

arXiv:2604.08540v1 Announce Type: cross Abstract: Text-to-Audio-Video (T2AV) generation is rapidly becoming a core interface for media creation, yet its evaluation remains fragmented. Existing benchma

Bayesian E(3)-Equivariant Interatomic Potential with Iterative Restratification of Many-body Message Passing

Model ReleasesDGX agent

arXiv:2510.03046v2 Announce Type: replace Abstract: Machine learning potentials (MLPs) have become essential for large-scale atomistic simulations, enabling ab initio-level accuracy with computational

Before We Trust Them: Decision-Making Failures in Navigation of Foundation Models

Model ReleasesDGX agent

arXiv:2601.05529v5 Announce Type: replace Abstract: High success rates on navigation-related tasks do not necessarily translate into reliable decision making by foundation models. To examine this gap,

BenchBrowser: Retrieving Evidence for Evaluating Benchmark Validity

Model ReleasesDGX agent

arXiv:2603.18019v2 Announce Type: replace Abstract: Do language model benchmarks actually measure what practitioners intend them to ? High-level metadata is too coarse to convey the granular reality o

Benchmarking LLM Tool-Use in the Wild

Model ReleasesDGX agent

arXiv:2604.06185v1 Announce Type: cross Abstract: Fulfilling user needs through Large Language Model multi-turn, multi-step tool-use is rarely a straightforward process. Real user interactions are inh

Beyond Case Law: Evaluating Structure-Aware Retrieval and Safety in Statute-Centric Legal QA

Model ReleasesDGX agent

arXiv:2604.06173v1 Announce Type: cross Abstract: Legal QA benchmarks have predominantly focused on case law, overlooking the unique challenges of statute-centric regulatory reasoning. In statutory do

Beyond Facts: Benchmarking Distributional Reading Comprehension in Large Language Models

Model ReleasesDGX agent

arXiv:2604.06201v1 Announce Type: cross Abstract: While most reading comprehension benchmarks for LLMs focus on factual information that can be answered by localizing specific textual evidence, many r

Beyond Mamba: Enhancing State-space Models with Deformable Dilated Convolutions for Multi-scale Traffic Object Detection

Model ReleasesDGX agent

arXiv:2604.08038v1 Announce Type: new Abstract: In a real-world traffic scenario, varying-scale objects are usually distributed in a cluttered background, which poses great challenges to accurate dete

Beyond Social Pressure: Benchmarking Epistemic Attack in Large Language Models

Model ReleasesDGX agent

arXiv:2604.07749v1 Announce Type: new Abstract: Large language models (LLMs) can shift their answers under pressure in ways that reflect accommodation rather than reasoning. Prior work on sycophancy h

Beyond Surface Artifacts: Capturing Shared Latent Forgery Knowledge Across Modalities

Model ReleasesDGX agent

arXiv:2604.07763v1 Announce Type: new Abstract: As generative artificial intelligence evolves, deepfake attacks have escalated from single-modality manipulations to complex, multimodal threats. Existi

BiScale-GTR: Fragment-Aware Graph Transformers for Multi-Scale Molecular Representation Learning

Model ReleasesDGX agent

arXiv:2604.06336v1 Announce Type: cross Abstract: Graph Transformers have recently attracted attention for molecular property prediction by combining the inductive biases of graph neural networks (GNN

Blending Human and LLM Expertise to Detect Hallucinations and Omissions in Mental Health Chatbot Responses

Model ReleasesDGX agent

arXiv:2604.06216v1 Announce Type: cross Abstract: As LLM-powered chatbots are increasingly deployed in mental health services, detecting hallucinations and omissions has become critical for user safet

Blind Refusal: Language Models Refuse to Help Users Evade Unjust, Absurd, and Illegitimate Rules

Model ReleasesDGX agent

arXiv:2604.06233v1 Announce Type: new Abstract: Safety-trained language models routinely refuse requests for help circumventing rules. But not all rules deserve compliance. When users ask for help eva

Bootstrapping Sign Language Annotations with Sign Language Models

Model ReleasesDGX agent

arXiv:2604.07606v1 Announce Type: new Abstract: AI-driven sign language interpretation is limited by a lack of high-quality annotated data. New datasets including ASL STEM Wiki and FLEURS-ASL contain

Break Me If You Can: Self-Jailbreaking of Aligned LLMs via Lexical Insertion Prompting

Model ReleasesDGX agent

arXiv:2601.02670v2 Announce Type: replace Abstract: We introduce self-jailbreaking, a threat model in which an aligned LLM guides its own compromise. Unlike most jailbreak techniques, which oft

BREAKING: GLM 5.1 by @Zai_org overwhelmingly dominates design-centric coding tasks amongst open-weight models In the categories featured bel…

Model ReleasesDGX agent

BREAKING: GLM 5.1 by @Zai_org overwhelmingly dominates design-centric coding tasks amongst open-weight models In the categories featured below, it is most comparable to Opus 4.6 by @AnthropicAI at ~1/

Bridging Theory and Practice in Crafting Robust Spiking Reservoirs

Model ReleasesDGX agent

arXiv:2604.06395v1 Announce Type: new Abstract: Spiking reservoir computing provides an energy-efficient approach to temporal processing, but reliably tuning reservoirs to operate at the edge-of-chaos

Broken by Default: A Formal Verification Study of Security Vulnerabilities in AI-Generated Code

Model ReleasesDGX agent

arXiv:2604.05292v2 Announce Type: replace-cross Abstract: AI coding assistants are now used to generate production code in security-sensitive domains, yet the exploitability of their outputs remains u

BTC-LLM: Efficient Sub-1-Bit LLM Quantization via Learnable Transformation and Binary Codebook

Model ReleasesDGX agent

arXiv:2506.12040v2 Announce Type: replace-cross Abstract: Binary quantization represents the most extreme form of compression, reducing weights to +/-1 for maximal memory and computational efficiency.

CADENCE: Context-Adaptive Depth Estimation for Navigation and Computational Efficiency

Model ReleasesDGX agent

arXiv:2604.07286v1 Announce Type: cross Abstract: Autonomous vehicles deployed in remote environments typically rely on embedded processors, compact batteries, and lightweight sensors. These hardware

Calibration of a neural network ocean closure for improved mean state and variability

Model ReleasesDGX agent

arXiv:2604.06398v1 Announce Type: cross Abstract: Global ocean models exhibit biases in the mean state and variability, particularly at coarse resolution, where mesoscale eddies are unresolved. To add

CAMO: A Class-Aware Minority-Optimized Ensemble for Robust Language Model Evaluation on Imbalanced Data

Model ReleasesDGX agent

arXiv:2604.07583v1 Announce Type: new Abstract: Real-world categorization is severely hampered by class imbalance because traditional ensembles favor majority classes, which lowers minority performanc

CAMotion: A High-Quality Benchmark for Camouflaged Moving Object Detection in the Wild

Model ReleasesDGX agent

arXiv:2604.08287v1 Announce Type: new Abstract: Discovering camouflaged objects is a challenging task in computer vision due to the high similarity between camouflaged objects and their surroundings.

Can Vision Language Models Judge Action Quality? An Empirical Evaluation

Model ReleasesDGX agent

arXiv:2604.08294v1 Announce Type: cross Abstract: Action Quality Assessment (AQA) has broad applications in physical therapy, sports coaching, and competitive judging. Although Vision Language Models

CASE: Cadence-Aware Set Encoding for Large-Scale Next Basket Repurchase Recommendation

Model ReleasesDGX agent

arXiv:2604.06718v2 Announce Type: cross Abstract: Repurchase behavior is a primary signal in large-scale retail recommendation, particularly in categories with frequent replenishment: many items in a

Chunks as Arms: Multi-Armed Bandit-Guided Sampling for Long-Context LLM Preference Optimization

Model ReleasesDGX agent

arXiv:2508.13993v2 Announce Type: replace Abstract: Long-context modeling is critical for a wide range of real-world tasks, including long-context question answering, summarization, and complex reason

Claude for Word is now in beta. Draft, edit, and revise documents directly from the sidebar. Claude preserves your formatting, and edits app…

Model ReleasesDGX agent

Claude for Word is now in beta. Draft, edit, and revise documents directly from the sidebar. Claude preserves your formatting, and edits appear as tracked changes. Available on Team and Enterprise pla

Claude now supports dynamic looping. If you run /loop without passing an interval, Claude will dynamically schedule the next tick based on y…

Model ReleasesDGX agent

Claude now supports dynamic looping. If you run /loop without passing an interval, Claude will dynamically schedule the next tick based on your task. It also may directly use the Monitor tool to bypas

ClawBench: Can AI Agents Complete Everyday Online Tasks?

Model ReleasesDGX agent

arXiv:2604.08523v1 Announce Type: new Abstract: AI agents may be able to automate your inbox, but can they automate other routine aspects of your life? Everyday online tasks offer a realistic yet unso

ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces

Model ReleasesDGX agent

arXiv:2604.05172v2 Announce Type: replace Abstract: Large language model (LLM) agents are increasingly deployed to automate productivity tasks (e.g., email, scheduling, document management), but evalu

Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding

Model ReleasesDGX agent

arXiv:2603.18472v2 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) perform strongly on natural images, yet their ability to understand discrete visual symbols remains u

Commander-GPT: Dividing and Routing for Multimodal Sarcasm Detection

Model ReleasesDGX agent

arXiv:2506.19420v2 Announce Type: replace Abstract: Multimodal sarcasm understanding is a high-order cognitive task. Although large language models (LLMs) have shown impressive performance on many dow

Consistency-Guided Decoding with Proof-Driven Disambiguation for Three-Way Logical Question Answering

Model ReleasesDGX agent

arXiv:2604.06196v1 Announce Type: cross Abstract: Three-way logical question answering (QA) assigns True/False/Unknown to a hypothesis H given a premise set S. While modern large language models

ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training

Model ReleasesDGX agent

arXiv:2604.07484v1 Announce Type: cross Abstract: Generative reward models (GRMs) have emerged as a promising approach for aligning Large Language Models (LLMs) with human preferences by offering grea

Contextual Earnings-22: A Speech Recognition Benchmark with Custom Vocabulary in the Wild

Model ReleasesDGX agent

arXiv:2604.07354v1 Announce Type: new Abstract: The accuracy frontier of speech-to-text systems has plateaued on academic benchmarks.1 In contrast, industrial benchmarks and adoption in high-stakes do

Continual Visual Anomaly Detection on the Edge: Benchmark and Efficient Solutions

Model ReleasesDGX agent

arXiv:2604.06435v1 Announce Type: cross Abstract: Visual Anomaly Detection (VAD) is a critical task for many applications including industrial inspection and healthcare. While VAD has been extensively

ConvoLearn: A Dataset for Fine-Tuning Dialogic AI Tutors

Model ReleasesDGX agent

arXiv:2601.08950v3 Announce Type: replace Abstract: Despite their growing adoption in education, LLMs remain misaligned with the core principle of effective tutoring: the dialogic construction of know

CoreWeave inks multiyear cloud deal with Anthropic

Model ReleasesDGX agent

CoreWeave Inc. today announced that it has won a multiyear contract to supply Anthropic PBC with cloud infrastructure. The company’s shares closed 11% higher on the news. The data center capacity comm

CrashSight: A Phase-Aware, Infrastructure-Centric Video Benchmark for Traffic Crash Scene Understanding and Reasoning

Model ReleasesDGX agent

arXiv:2604.08457v1 Announce Type: new Abstract: Cooperative autonomous driving requires traffic scene understanding from both vehicle and infrastructure perspectives. While vision-language models (VLM

Create Expert Content: Local Testing of a Multi-Agent System with Memory

Model ReleasesDGX agent

In support of our mission to accelerate the developer journey on Google Cloud, we built Dev Signal: a multi-agent system designed to transform raw community signals into reliable technical guidance by

Cron Jobs — AI That Works on a Schedule Tell Qwen Code 'check if tests pass every 30 minutes' and it sets up a cron job in your session. No …

Model ReleasesDGX agent

Cron Jobs — AI That Works on a Schedule Tell Qwen Code 'check if tests pass every 30 minutes' and it sets up a cron job in your session. No crontab editing, no scripts to write. Use /loop commend. Wor

Cross-Lingual Transfer and Parameter-Efficient Adaptation in the Turkic Language Family: A Theoretical Framework for Low-Resource Language Models

Model ReleasesDGX agent

arXiv:2604.06202v1 Announce Type: cross Abstract: Large language models (LLMs) have transformed natural language processing, yet their capabilities remain uneven across languages. Most multilingual mo

CryoSplat: Gaussian Splatting for Cryo-EM Homogeneous Reconstruction

Model ReleasesDGX agent

arXiv:2508.04929v4 Announce Type: replace-cross Abstract: As a critical modality for structural biology, cryogenic electron microscopy (cryo-EM) facilitates the determination of macromolecular structu

CycleChart: A Unified Consistency-Based Learning Framework for Bidirectional Chart Understanding and Generation

Model ReleasesDGX agent

arXiv:2512.19173v2 Announce Type: replace Abstract: Current chart-related tasks, such as chart generation (NL2Chart), chart schema parsing, chart data parsing, and chart question answering (ChartQA),

← Previous
1…360361362363364…369
Next →