AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,292 results
16 Apr 2026

Amazon launches its first smart warehouse in Shenzhen, aiming to cut local merchant storage costs by up to 45% as competition with Shein and Temu intensifies (Iris Deng/South China Morning Post)

Model ReleasesDGX agent

Iris Deng / South China Morning Post: Amazon launches its first smart warehouse in Shenzhen, aiming to cut local merchant storage costs by up to 45% as competition with Shein and Temu intensifies — Am

An Empirical Investigation of Practical LLM-as-a-Judge Improvement Techniques on RewardBench 2

Model ReleasesDGX agent

arXiv:2604.13717v1 Announce Type: new Abstract: LLM-as-a-judge, using a language model to score or rank candidate responses, is widely used as a scalable alternative to human evaluation in RLHF pipeli

An Optimal Transport-driven Approach for Cultivating Latent Space in Online Incremental Learning


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

arXiv:2211.16780v3 Announce Type: replace-cross Abstract: In online incremental learning, data continuously arrives with substantial distributional shifts, creating a significant challenge because pre

Analog Optical Inference on Million-Record Mortgage Data

Model ReleasesDGX agent

arXiv:2604.13251v1 Announce Type: new Abstract: Analog optical computers promise large efficiency gains for machine learning inference, yet no demonstration has moved beyond small-scale image benchmar

Anthropic launches Claude Opus 4.7 with coding, visual reasoning improvements

Model ReleasesDGX agent

Anthropic PBC today opened access to Claude Opus 4.7, the latest addition to its popular line of large language models. The company says that the LLM is significantly better than its predecessor at co

Anthropic releases a new Opus model amid Mythos Preview buzz

Model ReleasesDGX agent

Anthropic has released its most powerful 'generally available' model to date: Claude Opus 4.7. The company called it a step up from Opus 4.6 for advanced software engineering tasks, particularly in co

Anthropic says Opus 4.7 hits 80.6% on Document Reasoning — up from 57.1%. But 'reasoning about documents' ≠ 'parsing documents for agents.' …

Model ReleasesDGX agent

Anthropic says Opus 4.7 hits 80.6% on Document Reasoning — up from 57.1%. But 'reasoning about documents' ≠ 'parsing documents for agents.' We ran it on ParseBench. → Charts: 13.5% → 55.8% (+42.3) — h

As agentic AI overwhelms enterprise defenses, Oracle makes the case for security baked into the database

Model ReleasesDGX agent

As organizations transition from experimentation with AI to full-scale production, the demand for mission-critical security and absolute data availability has become the primary benchmark for enterpri

ASTER: Latent Pseudo-Anomaly Generation for Unsupervised Time-Series Anomaly Detection

Model ReleasesDGX agent

arXiv:2604.13924v1 Announce Type: cross Abstract: Time-series anomaly detection (TSAD) is critical in domains such as industrial monitoring, healthcare, and cybersecurity, but it remains challenging d

ASTRA: Enhancing Multi-Subject Generation with Retrieval-Augmented Pose Guidance and Disentangled Position Embedding

Model ReleasesDGX agent

arXiv:2604.13938v1 Announce Type: new Abstract: Subject-driven image generation has shown great success in creating personalized content, but its capabilities are largely confined to single subjects i

AudioX: A Unified Framework for Anything-to-Audio Generation

Model ReleasesDGX agent

arXiv:2503.10522v4 Announce Type: replace-cross Abstract: Audio and music generation based on flexible multimodal control signals is a widely applicable topic, with the following key challenges: 1) a

Auto-FP: An Experimental Study of Automated Feature Preprocessing for Tabular Data

Model ReleasesDGX agent

arXiv:2310.02540v2 Announce Type: replace Abstract: Classical machine learning models, such as linear models and tree-based models, are widely used in industry. These models are sensitive to data dist

Autonomous Multi-objective Alloy Design through Simulation-guided Optimization

Model ReleasesDGX agent

arXiv:2507.16005v2 Announce Type: replace-cross Abstract: Alloy discovery is constrained by vast compositional spaces, competing objectives, and prohibitive experimental costs. Although simulations an

BenGER: A Collaborative Web Platform for End-to-End Benchmarking of German Legal Tasks

Model ReleasesDGX agent

arXiv:2604.13583v1 Announce Type: new Abstract: Evaluating large language models (LLMs) for legal reasoning requires workflows that span task design, expert annotation, model execution, and metric-bas

Beyond Static Personas: Situational Personality Steering for Large Language Models

Model ReleasesDGX agent

arXiv:2604.13846v1 Announce Type: new Abstract: Personalized Large Language Models (LLMs) facilitate more natural, human-like interactions in human-centric applications. However, existing personalizat

Beyond Uniform Sampling: Synergistic Active Learning and Input Denoising for Robust Neural Operators

Model ReleasesDGX agent

arXiv:2604.13316v1 Announce Type: new Abstract: Neural operators have emerged as fast surrogate models for physics simulations, yet they remain acutely vulnerable to adversarial perturbations, a criti

Breaking the Generator Barrier: Disentangled Representation for Generalizable AI-Text Detection

Model ReleasesDGX agent

arXiv:2604.13692v1 Announce Type: new Abstract: As large language models (LLMs) generate text that increasingly resembles human writing, the subtle cues that distinguish AI-generated content from huma

Bridging MARL to SARL: An Order-Independent Multi-Agent Transformer via Latent Consensus

Model ReleasesDGX agent

arXiv:2604.13472v1 Announce Type: new Abstract: Cooperative multi-agent reinforcement learning (MARL) is widely used to address large joint observation and action spaces by decomposing a centralized c

Built an political benchmark for LLMs. KIMI K2 can't answer about Taiwan (Obviously). GPT-5.3 refuses 100% of questions when given an opt-out. [P]

Model ReleasesDGX agent

A researcher on r/MachineLearning built a political benchmark to evaluate how various LLMs handle sensitive geopolitical and politically contentious questions. Key findings include that Kimi K2 (Moons

Can Large Language Models Reliably Extract Physiology Index Values from Coronary Angiography Reports?

Model ReleasesDGX agent

arXiv:2604.13077v1 Announce Type: new Abstract: Coronary angiography (CAG) reports contain clinically relevant physiological measurements, yet this information is typically in the form of unstructured

CANVAS: Continuity-Aware Narratives via Visual Agentic Storyboarding

Model ReleasesDGX agent

arXiv:2604.13452v1 Announce Type: new Abstract: Long-form visual storytelling requires maintaining continuity across shots, including consistent characters, stable environments, and smooth scene trans

Chain of Uncertain Rewards with Large Language Models for Reinforcement Learning

Model ReleasesDGX agent

arXiv:2604.13504v1 Announce Type: cross Abstract: Designing effective reward functions is a cornerstone of reinforcement learning (RL), yet it remains a challenging and labor-intensive process due to

Claude Opus 4.7 is now available as an Agent Preview inside of Devin! Anthropic has clearly optimized Claude Opus 4.7 for long-horizon auton…

Model ReleasesDGX agent

Claude Opus 4.7 is now available as an Agent Preview inside of Devin! Anthropic has clearly optimized Claude Opus 4.7 for long-horizon autonomy, unlocking a class of deep investigation work we couldn'

Claude Opus 4.7 is now available in Cursor. We've found it to be impressively autonomous and more creative in its reasoning. We're launching…

Model ReleasesDGX agent

Cursor has integrated Claude Opus 4.7 into its platform, highlighting the model's impressive autonomous capabilities and enhanced creative reasoning abilities. The announcement suggests a new feature

Claude Opus 4.7 is now available in Windsurf 2.0! Anthropic has clearly optimized Claude Opus 4.7 for sustained reasoning over long runs. Ag…

Model ReleasesDGX agent

Claude Opus 4.7 is now available in Windsurf 2.0! Anthropic has clearly optimized Claude Opus 4.7 for sustained reasoning over long runs. Agents stay on track longer without intervention, so engineers

Claude Opus 4.7 is now the default orchestration model powering Computer. It's also available for Max subscribers on Perplexity web, iOS, an…

Model ReleasesDGX agent

Claude Opus 4.7 has been set as the default orchestration model for Anthropic's Computer product. The model is also available to Max subscribers on Perplexity's web platform and iOS application.

Claude remains irreducibly Claude. If you know, you know. (The fact that models have distinct personalities that are consistent across gener…

Model ReleasesDGX agent

Claude remains irreducibly Claude. If you know, you know. (The fact that models have distinct personalities that are consistent across generations is technically interesting, it also makes it very eas

CodeFlowBench: A Multi-turn, Iterative Benchmark for Complex Code Generation

Model ReleasesDGX agent

arXiv:2504.21751v4 Announce Type: replace-cross Abstract: Modern software development demands code that is maintainable, testable, and scalable by organizing the implementation into modular components

Codex for (almost) everything

Model ReleasesDGX agent

OpenAI's Codex is a large language model trained on publicly available code from the internet that can understand and generate code in dozens of programming languages. It powers GitHub Copilot and can

Coherence in the brain unfolds across separable temporal regimes

Model ReleasesDGX agent

arXiv:2512.20481v4 Announce Type: replace-cross Abstract: To maintain coherence in language, the brain must satisfy key competing temporal demands: the gradual accumulation of meaning across extended

CollabCoder: Plan-Code Co-Evolution via Collaborative Decision-Making for Efficient Code Generation

Model ReleasesDGX agent

arXiv:2604.13946v1 Announce Type: cross Abstract: Automated code generation remains a persistent challenge in software engineering, as conventional multi-agent frameworks are often constrained by stat

Common to Whom? Regional Cultural Commonsense and LLM Bias in India

Model ReleasesDGX agent

arXiv:2601.15550v3 Announce Type: replace Abstract: Existing cultural commonsense benchmarks treat nations as monolithic, assuming uniform practices within national boundaries. But does cultural commo

Correct Chains, Wrong Answers: Dissociating Reasoning from Output in LLM Logic

Model ReleasesDGX agent

arXiv:2604.13065v1 Announce Type: new Abstract: LLMs can execute every step of chain-of-thought reasoning correctly and still produce wrong final answers. We introduce the Novel Operator Test, a bench

Correct Prediction, Wrong Steps? Consensus Reasoning Knowledge Graph for Robust Chain-of-Thought Synthesis

Model ReleasesDGX agent

arXiv:2604.14121v1 Announce Type: new Abstract: LLM reasoning traces suffer from complex flaws -- *Step Internal Flaws* (logical errors, hallucinations, etc.) and *Step-wise Flaws* (overthinking, unde

Counterfactual Peptide Editing for Causal TCR--pMHC Binding Inference

Model ReleasesDGX agent

arXiv:2604.13256v1 Announce Type: new Abstract: Neural models for TCR-pMHC binding prediction are susceptible to shortcut learning: they exploit spurious correlations in training data -- such as pepti

Covariance-adapting algorithm for semi-bandits with application to sparse rewards

Model ReleasesDGX agent

arXiv:2604.13738v1 Announce Type: cross Abstract: We investigate stochastic combinatorial semi-bandits, where the entire joint distribution of outcomes impacts the complexity of the problem instance (

Cracking the Code of Juxtaposition: Can AI Models Understand the Humorous Contradictions

Model ReleasesDGX agent

arXiv:2405.19088v3 Announce Type: replace Abstract: Recent advancements in large multimodal language models have demonstrated remarkable proficiency across a wide range of tasks. Yet, these models sti

Data-driven Learning of Probabilistic Model of Binary Droplet Collision for Spray Simulation

Model ReleasesDGX agent

arXiv:2604.13594v1 Announce Type: cross Abstract: Binary droplet collisions are ubiquitous in dense sprays. Traditional deterministic models cannot adequately represent transitional and stochastic beh

Databricks on Google Cloud: Innovate Faster. Smarter. Together.

Model ReleasesDGX agent

Databricks and Google Cloud have partnered to enable organizations to build and deploy data and AI solutions more efficiently. The collaboration integrates Databricks' lakehouse platform with Google C

datasette.io news preview

Model ReleasesDGX agent

Tool: datasette.io news preview The datasette.io website has a news section built from this news.yaml file in the underlying GitHub repository. The YAML format looks like this: - date: 2026-04-15 body

Decoding the Delta: Unifying Remote Sensing Change Detection and Understanding with Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2604.14044v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) excel in general vision-language tasks, their application to remote sensing change understanding is hinde

DeepL, best known for its text translation tools, launches DeepL Voice-to-Voice, which enables real-time spoken translation, with add-ons for services like Zoom (Ivan Mehta/TechCrunch)

Model ReleasesDGX agent

Ivan Mehta / TechCrunch: DeepL, best known for its text translation tools, launches DeepL Voice-to-Voice, which enables real-time spoken translation, with add-ons for services like Zoom — DeepL, a tra

DeEscalWild: A Real-World Benchmark for Automated De-Escalation Training with SLMs

Model ReleasesDGX agent

arXiv:2604.13075v1 Announce Type: new Abstract: Effective de-escalation is critical for law enforcement safety and community trust, yet traditional training methods lack scalability and realism. While

Defending Your Enterprise When AI Models Can Find Vulnerabilities Faster Than Ever

Model ReleasesDGX agent

Introduction Advances in AI model-powered exploitation have demonstrated that general-purpose AI models can excel at vulnerability discovery, even without being purpose-built for the task. Eventually,

Dental-TriageBench: Benchmarking Multimodal Reasoning for Hierarchical Dental Triage

Model ReleasesDGX agent

arXiv:2604.13060v1 Announce Type: new Abstract: Dental triage is a safety-critical clinical routing task that requires integrating multimodal clinical information (e.g., patient complaints and radiogr

Design Space Exploration of Hybrid Quantum Neural Networks for Chronic Kidney Disease

Model ReleasesDGX agent

arXiv:2604.13608v1 Announce Type: new Abstract: Hybrid Quantum Neural Networks (HQNNs) have recently emerged as a promising paradigm for near-term quantum machine learning. However, their practical pe

DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis

Model ReleasesDGX agent

arXiv:2604.13416v1 Announce Type: new Abstract: Advances in radiance fields have enabled photorealistic novel view synthesis. In several domains, large-scale real-world datasets have been developed to

Do We Still Need Humans in the Loop? Comparing Human and LLM Annotation in Active Learning for Hostility Detection

Model ReleasesDGX agent

arXiv:2604.13899v1 Announce Type: new Abstract: Instruction-tuned LLMs can annotate thousands of instances from a short prompt at negligible cost. This raises two questions for active learning (AL): c

Document-tuning for robust alignment to animals

Model ReleasesDGX agent

arXiv:2604.13076v1 Announce Type: new Abstract: We investigate the robustness of value alignment via finetuning with synthetic documents, using animal compassion as a value that is both important in i

Drowsiness-Aware Adaptive Autonomous Braking System based on Deep Reinforcement Learning for Enhanced Road Safety

Model ReleasesDGX agent

arXiv:2604.13878v1 Announce Type: new Abstract: Driver drowsiness significantly impairs the ability to accurately judge safe braking distances and is estimated to contribute to 10%-20% of road acciden

Efficient Multi-View 3D Object Detection by Dynamic Token Selection and Fine-Tuning

Model ReleasesDGX agent

arXiv:2604.13586v1 Announce Type: new Abstract: Existing multi-view three-dimensional (3D) object detection approaches widely adopt large-scale pre-trained vision transformer (ViT)-based foundation mo

EmbodiedClaw: Conversational Workflow Execution for Embodied AI Development

Model ReleasesDGX agent

arXiv:2604.13800v1 Announce Type: new Abstract: Embodied AI research is increasingly moving beyond single-task, single-environment policy learning toward multi-task, multi-scene, and multi-model setti

EMGFlow: Robust and Efficient Surface Electromyography Synthesis via Flow Matching

Model ReleasesDGX agent

arXiv:2604.13685v1 Announce Type: cross Abstract: Deep learning-based surface electromyography (sEMG) gesture recognition is frequently bottlenecked by data scarcity and limited subject diversity. Whi

Enhancing Confidence Estimation in Telco LLMs via Twin-Pass CoT-Ensembling

Model ReleasesDGX agent

arXiv:2604.13271v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly applied to complex telecommunications tasks, including 3GPP specification analysis and O-RAN network troub

ESCAPE: Episodic Spatial Memory and Adaptive Execution Policy for Long-Horizon Mobile Manipulation

Model ReleasesDGX agent

arXiv:2604.13633v1 Announce Type: new Abstract: Coordinating navigation and manipulation with robust performance is essential for embodied AI in complex indoor environments. However, as tasks extend o

Evaluating LLM-Based Translation of a Low-Resource Technical Language: The Medical and Philosophical Greek of Galen

Model ReleasesDGX agent

arXiv:2602.24119v2 Announce Type: replace Abstract: Purpose: This study evaluates the quality of commercial large language model (LLM) machine translation (MT) for Ancient Greek technical prose and be

Evaluating Supervised Machine Learning Models: Principles, Pitfalls, and Metric Selection

Model ReleasesDGX agent

arXiv:2604.13882v1 Announce Type: new Abstract: The evaluation of supervised machine learning models is a critical stage in the development of reliable predictive systems. Despite the widespread avail

Evaluating the Evaluator: Problems with SemEval-2020 Task 1 for Lexical Semantic Change Detection

Model ReleasesDGX agent

arXiv:2604.13232v1 Announce Type: new Abstract: This discussion paper re-examines SemEval-2020 Task 1, the most influential shared benchmark for lexical semantic change detection, through a three-part

Evaluating the Formal Reasoning Capabilities of Large Language Models through Chomsky Hierarchy

Model ReleasesDGX agent

arXiv:2604.02709v2 Announce Type: replace Abstract: The formal reasoning capabilities of LLMs are crucial for advancing automated software engineering. However, existing benchmarks for LLMs lack syste

EVE: A Domain-Specific LLM Framework for Earth Intelligence

Model ReleasesDGX agent

arXiv:2604.13071v1 Announce Type: new Abstract: We introduce Earth Virtual Expert (EVE), the first open-source, end-to-end initiative for developing and deploying domain-specialized LLMs for Earth Int

← Previous
1…340341342343344…372
Next →