AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,619Total entries
1Added by human
84,618Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,603 results
10 Jun 2026

Benchmarking Knowledge Editing using Logical Rules

Model ReleasesDGX agent

arXiv:2606.10554v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly deployed in real-world applications that require access to up-to-date knowledge. However, retraining LLM

Benchmarking stereo reconstruction for 3D printable Martian terrain models

Model ReleasesDGX agent

arXiv:2606.10364v1 Announce Type: new Abstract: Reconstructing printable 3D models from Mars rover imagery is challenging because Martian terrain is low-texture, irregular, and partially observed. We

BenSyc: Benchmarking Conversational Sycophancy and Human Alignment in LLMs for Bengali Contexts

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.10061v1 Announce Type: new Abstract: Large language models (LLMs) increasingly participate in emotionally sensitive social conversations, where responses may shift from balanced support tow

Beyond APIs: Probing the Limits of MLLMs in Physical Tool Use

Model ReleasesDGX agent

arXiv:2606.10803v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) excel at utilizing digital APIs and increasingly serve as the 'brain' of embodied AI, instructing robots to i

Beyond Memorization: Distinguishing Between Pattern-Based and Epistemic Reasoning in LLMs Using Epistemic Puzzles

Model ReleasesDGX agent

arXiv:2603.21350v2 Announce Type: replace Abstract: Epistemic reasoning requires agents to infer the state of the world from partial observations and information about other agents' knowledge. Prior w

Beyond Model Size: Probing the Gaps in Visual in-Context Learning by Training a Tiny Model

Model ReleasesDGX agent

arXiv:2606.10905v1 Announce Type: new Abstract: Visual in-Context Learning (VICL) aims at making progress towards adaptive vision models, that can -- based on a few examples -- adapt to a new task at

Bittensor Agent Arenas as a Trajectory Primitive: Distilling a Shopping Agent from ShoppingBench Subnet Traces

Model ReleasesDGX agent

arXiv:2606.10064v1 Announce Type: cross Abstract: Small-model agentic post-training is bottlenecked less by the algorithm than by the trajectory substrate it consumes. Leading recipes (RLVR, group-rel

Boosting Graph Robustness Against Backdoor Attacks: An Over-Similarity Perspective

Model ReleasesDGX agent

arXiv:2502.01272v3 Announce Type: replace Abstract: Graph Neural Networks (GNNs) have achieved notable success in tasks such as social and transportation networks. However, recent studies have highlig

CableRobotGraphSim: A Graph Neural Network for Modeling Partially Observable Cable-Driven Robot Dynamics

Model ReleasesDGX agent

arXiv:2602.21331v2 Announce Type: replace Abstract: General-purpose simulators have accelerated the development of robots. Traditional simulators based on first-principles, however, typically require

Can Image Models Imagine Time? ImageTime: A Novel Benchmark for Probing Visual World Modeling Through Spatiotemporal Consistency

Model ReleasesDGX agent

arXiv:2606.10620v1 Announce Type: cross Abstract: Image generation models now produce high-quality static images, yet their ability to represent how a visual world changes over time remains poorly und

Can Multi-Agent LLMs Identify Their Peers? Stylometric Fingerprinting in Role-Constrained Political Analysis

Model ReleasesDGX agent

arXiv:2606.09854v1 Announce Type: cross Abstract: Multi-agent large language model (LLM) pipelines for political statement analysis are vulnerable to peer-preservation bias: models tend to protect pee

ChartLens: A Dual-Branch Framework for Chart Data Correction and Factual Summary Refinement

Model ReleasesDGX agent

arXiv:2606.10640v1 Announce Type: new Abstract: In this report, we present our champion solution for the DataMFM Challenge Track 2: Chart Understanding. This track requires models to recover structure

CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs

Model ReleasesDGX agent

arXiv:2606.11063v1 Announce Type: new Abstract: AI control protocols oversee untrusted models by monitoring their actions and modifying potentially unsafe steps, often using a trusted model. This part

CITRAS-FM: Tiny Time Series Foundation Model for Covariate-Informed Zero-Shot Forecasting

Model ReleasesDGX agent

arXiv:2606.10798v1 Announce Type: new Abstract: Pretrained time series foundation models (TSFMs) have enabled zero-shot forecasting on unseen target series. However, existing TSFMs often incur high co

Claude Fable 5 is now available in Computer as an orchestrator model. This is Anthropic's state-of-the-art model for long, complex tasks. Av…

Model ReleasesDGX agent

Claude Fable 5 is now available in Computer as an orchestrator model designed by Anthropic for handling long and complex tasks. This represents Anthropic's latest state-of-the-art model offering enhan

Claude Fable 5 thinks document parsing is beneath it It is absolutely crushing on all reasoning-intensive/long horizon benchmarks: SWE-Bench…

Model ReleasesDGX agent

Claude Fable 5 thinks document parsing is beneath it It is absolutely crushing on all reasoning-intensive/long horizon benchmarks: SWE-Bench Pro, FrontierCode, GDPval, Runescape, etc. But for document

Claude Fable won’t answer basic biology questions

Model ReleasesDGX agent

Anthropic just released Claude Fable 5, calling it the most powerful AI model it has ever made widely available and praising its skills in biology, among others. But the model won't answer basic biolo

CleanPatrick: A Benchmark for Image Data Cleaning

Model ReleasesDGX agent

arXiv:2505.11034v2 Announce Type: replace-cross Abstract: Robust machine learning depends on clean data, yet current image data cleaning benchmarks rely on synthetic noise or narrow human studies, lim

CLP: Collocation-Length Prediction for Zero-Loss Adaptive Multi-Token Inference

Model ReleasesDGX agent

arXiv:2606.10935v1 Announce Type: cross Abstract: Large language model inference is bottlenecked by autoregressive decoding, where each token requires a full forward pass. Multi-token prediction (MTP)

ClusBench: The Clustering Benchmark Data Resource You've All Been Waiting For (?)

Model ReleasesDGX agent

arXiv:2606.10673v1 Announce Type: cross Abstract: Although some very common test beds exist for assessing the performance of clustering methods, large scale benchmarking is typically limited to relati

CoCoSI: Collaborative Cognitive Map Construction for Spatial Intelligence

Model ReleasesDGX agent

arXiv:2606.10401v1 Announce Type: new Abstract: Spatial intelligence is a key frontier for multimodal large language models (MLLMs), enabling them to reason about the physical world from visual experi

CodeAlchemy: Synthetic Code Rewriting at Scale

Model ReleasesDGX agent

arXiv:2606.10087v1 Announce Type: new Abstract: Pre-training on raw code teaches syntax but provides sparse signal for diverse real-world task formats. While synthetic data has proven transformative f

Cohere Transcribe, our open-source speech recognition model, is #1 on the new @huggingface Far-Field ASR benchmark.

Model ReleasesDGX agent

Cohere has released Transcribe, an open-source speech recognition model that achieved the top ranking on Hugging Face's newly established Far-Field Automatic Speech Recognition (ASR) benchmark. The mo

CollabSkill: Evaluating Human-Agent Collaboration On Real-World Tasks

Model ReleasesDGX agent

arXiv:2606.09833v1 Announce Type: cross Abstract: AI agents are reshaping the workspace, leading to drastic change of how humans work. Despite the considerable potential of human-agent collaboration b

ComBench: A Benchmark for Rigorous Proof Reasoning and Constructive Realization in Olympiad-Level Combinatorics

Model ReleasesDGX agent

arXiv:2606.10479v1 Announce Type: new Abstract: Combinatorics is central to Olympiad-level mathematical problem solving, requiring deep discrete reasoning, creative constructions, and rigorous structu

CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data

Model ReleasesDGX agent

arXiv:2601.18026v2 Announce Type: replace Abstract: Language identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, e

Compile Once, Differentiate Everywhere: A Differentiable Meta-Circular Interpreter

Model ReleasesDGX agent

arXiv:2606.09930v1 Announce Type: cross Abstract: The boundary between program execution and gradient-based optimization has long limited the use of code itself as a learnable scientific model. We pre

Constructing coherent spatial memory in LLM agents through graph rectification

Model ReleasesDGX agent

arXiv:2510.04195v2 Announce Type: replace Abstract: Given a map description through global traversal navigation instructions, an LLM can often infer the implicit spatial layout and answer user queries

Continual LLM Upcycling: A Predictor-Gated Bank-Wise Sparsity Training Recipe for Dense-to-Sparse LLMs

Model ReleasesDGX agent

arXiv:2606.10722v1 Announce Type: new Abstract: We study dense-to-sparse continual training as a way to construct channel-sparse large language models from dense checkpoints. Starting from a Qwen2.5-8

Continuous Neural Reparameterization as a Deep Geometric Prior for Robust Fixed-Chart UV Repair

Model ReleasesDGX agent

arXiv:2606.10050v1 Announce Type: cross Abstract: Traditional UV unwrapping relies on direct optimization of geometric distortion energies and can fail through invalid initialization, local minima, or

ConvMemory v2: A Recall-Preserving Top-10 Evidence Reranker for Conversational Memory Retrieval

Model ReleasesDGX agent

arXiv:2606.10842v1 Announce Type: new Abstract: We describe ConvMemory v2, an opt-in token-evidence reranker that sits after the lightweight ConvMemory v1 reranker and reorders only v1's protected top

CoTAL: Human-in-the-Loop Prompt Engineering for Generalizable Formative Assessment Scoring and Feedback

Model ReleasesDGX agent

arXiv:2504.02323v4 Announce Type: replace Abstract: Large language models (LLMs) have created new opportunities to assist teachers and support student learning. While researchers have explored various

Cybersecurity researchers complain that Claude Fable's guardrails are too strict, rejecting 'innocuous tasks' like reading blog posts or performing code reviews (Lorenzo Franceschi-Bicchierai/TechCrunch)

Model ReleasesDGX agent

Lorenzo Franceschi-Bicchierai / TechCrunch: Cybersecurity researchers complain that Claude Fable's guardrails are too strict, rejecting “innocuous tasks” like reading blog posts or performing code rev

Cyst-X: A Multi-Center MRI Benchmark and Federated Learning Framework for Malignancy-Risk Stratification of Pancreatic Cystic Neoplasm

Model ReleasesDGX agent

arXiv:2507.22017v4 Announce Type: replace-cross Abstract: Pancreatic cancer is projected to be the second-deadliest cancer by 2030, making early detection critical. Intraductal papillary mucinous neop

Dario Amodei says he doesn't know what role Claude played in a missile strike on an Iranian school, and its use in this instance didn't violate Anthropic's ToS (Bloomberg)

Model ReleasesDGX agent

Bloomberg: Dario Amodei says he doesn't know what role Claude played in a missile strike on an Iranian school, and its use in this instance didn't violate Anthropic's ToS — Anthropic PBC's boss said h

Data-Driven Dynamic Assortment in Online Platforms: Learning about Two Sides

Model ReleasesDGX agent

arXiv:2606.11118v1 Announce Type: new Abstract: We study a dynamic assortment problem on a two-sided service platform with incomplete information and heterogeneous customers in a discrete-time setting

Data-Driven Runway and Taxiway Exits Prediction of Landing Aircraft: A Case Study at Hartsfield-Jackson Atlanta International Airport

Model ReleasesDGX agent

arXiv:2606.11017v1 Announce Type: new Abstract: Airport surface operations increasingly constrain performance at high-throughput hubs. This study examines arrival taxi-in decisions at Hartsfield-Jacks

datasette-agent 0.2a0

Model ReleasesDGX agent

Release: datasette-agent 0.2a0 Highlights from the release notes: Tools can now ask the user questions mid-execution. Tools that declare a context parameter receive a ToolContext object, and await con

Day 0 Anthropic Fable 5 in ParseBench: We tested the model's advancements when it comes to document understanding. The model clearly peaks w…

Model ReleasesDGX agent

Day 0 Anthropic Fable 5 in ParseBench: We tested the model's advancements when it comes to document understanding. The model clearly peaks when it comes to adherence to the original text: 📃 Content fa

DB-3DME: From Dataset to Benchmark for Human-aligned Automatic 3D Mesh Evaluation

Model ReleasesDGX agent

arXiv:2606.10142v1 Announce Type: new Abstract: Recent advances in 3D generation have led to substantial improvements in realism, controllability, and efficiency, yet the evaluation of 3D assets remai

Deployment-Time Memorization in Foundation-Model Agents

Model ReleasesDGX agent

arXiv:2606.10062v1 Announce Type: new Abstract: Foundation-model agents are increasingly long-lived systems that remember users across interactions, making memorization an explicit deployment-time fun

Detecting Knowledge Gaps from Conversational AI Interactions Using Curriculum Prerequisite Graphs

Model ReleasesDGX agent

arXiv:2606.10736v1 Announce Type: cross Abstract: Large online courses generate thousands of student questions directed at conversational AI teaching assistants, yet these interaction logs remain larg

DiffusionGemma

Model ReleasesDGX agent

DiffusionGemma Last May Google briefly released an experimental Gemini Diffusion model. I tried the preview at the time and recorded it running at 857 tokens/second. It was an exciting model, but Goog

DiffusionGemma: 4x faster text generation

Model ReleasesDGX agent

DiffusionGemma is an experimental open model from Google DeepMind that uses text diffusion for exceptionally fast generation, moving beyond sequential token-by-token processing to generate entire bloc

Discovering Interpretable Multi-Parameter Control Policies for Evolutionary Algorithms Using Deep Reinforcement Learning

Model ReleasesDGX agent

arXiv:2606.10129v1 Announce Type: new Abstract: While deep Reinforcement Learning (deep-RL) has been increasingly applied to parameter control in evolutionary algorithms, rigorous theoretical analysis

Disjoint or Overlapping? Inference Windowing for Reconstruction-Based Time Series Anomaly Detection

Model ReleasesDGX agent

arXiv:2606.09874v1 Announce Type: new Abstract: Reconstruction-based methods are widely used for time series anomaly detection, where models are trained to reconstruct subsequences, and anomalies are

Divide-and-Conquer Modeling for the CTF-4-Science Lorenz Benchmark

Model ReleasesDGX agent

arXiv:2606.10084v1 Announce Type: cross Abstract: This work presents a divide-and-conquer modeling strategy for the CTF-4-Science Lorenz benchmark, which evaluates chaotic-system prediction across twe

Divide and Cooperate: Role-Decomposed Multi-Agent LLM Training with Cross-Agent Learning Signals

Model ReleasesDGX agent

arXiv:2606.10684v1 Announce Type: cross Abstract: Modern language agents which perform multi-step reasoning have shown strong performance in knowledge-intensive question answering. However, existing a

Do Vision-Language Models See or Guess? Measuring and Reducing Textual-Prior Reliance with a Phrasing-Controlled Benchmark

Model ReleasesDGX agent

arXiv:2606.10400v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly deployed where answers must follow from what is in the image, yet they often answer from textual priors,

Do VLMs Reason Like Engineers? A Benchmark and a Stage-wise Evaluation

Model ReleasesDGX agent

arXiv:2606.10833v1 Announce Type: new Abstract: Vision-Language Models (VLMs) demonstrate strong performance on general multimodal reasoning benchmarks, yet their ability to perform engineering reason

Domain Adapted Large Language Models for Additive Manufacturing

Model ReleasesDGX agent

arXiv:2603.22017v2 Announce Type: replace Abstract: This work presents a collection of multi-modal domain adapted large language models built upon the instruction tuned variants of open weight models

Don't waste SAM

Model ReleasesDGX agent

arXiv:2606.10696v1 Announce Type: new Abstract: Meta AI has recently released the Segment Anything Model (SAM), which demonstrates exceptional zero-shot image segmentation performance across various t

Dual-Branch Gated Fusion for Open-Set Audio Deepfake Source Tracing

Model ReleasesDGX agent

arXiv:2606.10223v1 Announce Type: cross Abstract: Attributing a synthetic utterance to its originating system remains an open challenge: closed-set models fail to reject unseen synthesizers and produc

Earth-OneVision: Extending Remote Sensing Multimodal Large Language Models to More Sensor Modalities and Tasks

Model ReleasesDGX agent

arXiv:2606.10819v1 Announce Type: cross Abstract: RS-MLLMs enable natural-language understanding and spatial reasoning over earth observation imagery. However, existing models support only a narrow ra

EEVEE: Towards Test-time Prompt Learning in the Real World for Self-Improving Agents

Model ReleasesDGX agent

arXiv:2606.11182v1 Announce Type: cross Abstract: In this paper, we propose EEVEE, the first multi-dataset test-time prompt learning framework for LLM agents, enabling test-time prompt learning under

Effective Reinforcement Learning for Agentic Search by Recycling Zero-Variance Queries During Training

Model ReleasesDGX agent

arXiv:2606.10709v1 Announce Type: cross Abstract: The use of GRPO-style algorithms has become the standard strategy for training LLM search agents under outcome-only rewards. With these algorithms, a

Effective Training Principles of Physical Reservoirs

Model ReleasesDGX agent

arXiv:2606.10130v1 Announce Type: cross Abstract: Reservoir computers benefit from the inherent complexity of optical phenomena, which provide rich, often nonlinear dynamics. However, training directl

Efficient RWKV-based Representation Learning for 3D Point Clouds

Model ReleasesDGX agent

arXiv:2606.10395v1 Announce Type: new Abstract: The recent receptance weighted key value (RWKV) model combines RNN-style recurrence, offering a linear-complexity alternative to Transformers' quadratic

Efficient-WAM: A 1B-Parameter World-Action Model with Low-Cost Future Imagination

Model ReleasesDGX agent

arXiv:2606.10040v1 Announce Type: new Abstract: World-Action Models (WAMs) have emerged as a promising paradigm for embodied control by coupling future visual prediction with action generation. Howeve

Enabling Progressive Whole-slide Image Analysis with Multi-scale Pyramidal Network

Model ReleasesDGX agent

arXiv:2602.01951v2 Announce Type: replace Abstract: Multiple-instance Learning (MIL) is commonly used for computational pathology (CPath), where multi-scale features are essential for capturing both f

← Previous
1…150151152153154…377
Next →