AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
All
84,661Total entries
1Added by human
84,660Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,620 results
Model Releases

Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It

DGX agent

arXiv:2606.11052v1 Announce Type: new Abstract: Chain-of-thought (CoT) supervised fine-tuning (SFT) is widely adopted to improve reasoning ability, yet we find that it systematically degrades long-con

model-releasesarxiv-cs-cl
10 Jun 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Attention Expansion: Enhancing Keyphrase Extraction from Long Documents with Attention-Augmented Contextualized Embeddings

DGX agent

arXiv:2606.10716v1 Announce Type: cross Abstract: Pre-trained language models (PLMs) have achieved strong performance in keyphrase extraction (KPE), largely due to their ability to generate rich conte

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

au-Rec: A Verifiable Benchmark for Agentic Recommender Systems

DGX agent

arXiv:2606.10156v1 Announce Type: cross Abstract: As recommender systems transition toward agentic, multi-turn conversational interfaces, evaluation paradigms have struggled to keep pace. Current benc

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

BadRobot: Jailbreaking Embodied LLM Agents in the Physical World

DGX agent

arXiv:2407.20242v5 Announce Type: replace-cross Abstract: Embodied AI represents systems where AI is integrated into physical entities. Large Language Model (LLM), which exhibits powerful language und

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

Benchmarking and Exploring the Capabilities of LLMs for Attack Investigations

DGX agent

arXiv:2606.10281v1 Announce Type: cross Abstract: This paper presents AuditBench, a new benchmark dataset for evaluating the capabilities of LLMs at investigating security-related system audit logs. W

model-releasesarxiv-cs-cl
10 Jun 2026
Model Releases

Benchmarking Knowledge Editing using Logical Rules

DGX agent

arXiv:2606.10554v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly deployed in real-world applications that require access to up-to-date knowledge. However, retraining LLM

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

Benchmarking stereo reconstruction for 3D printable Martian terrain models

DGX agent

arXiv:2606.10364v1 Announce Type: new Abstract: Reconstructing printable 3D models from Mars rover imagery is challenging because Martian terrain is low-texture, irregular, and partially observed. We

model-releasesarxiv-cs-cv
10 Jun 2026
Model Releases

BenSyc: Benchmarking Conversational Sycophancy and Human Alignment in LLMs for Bengali Contexts

DGX agent

arXiv:2606.10061v1 Announce Type: new Abstract: Large language models (LLMs) increasingly participate in emotionally sensitive social conversations, where responses may shift from balanced support tow

model-releasesarxiv-cs-cl
10 Jun 2026
Model Releases

Beyond APIs: Probing the Limits of MLLMs in Physical Tool Use

DGX agent

arXiv:2606.10803v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) excel at utilizing digital APIs and increasingly serve as the 'brain' of embodied AI, instructing robots to i

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

Beyond Memorization: Distinguishing Between Pattern-Based and Epistemic Reasoning in LLMs Using Epistemic Puzzles

DGX agent

arXiv:2603.21350v2 Announce Type: replace Abstract: Epistemic reasoning requires agents to infer the state of the world from partial observations and information about other agents' knowledge. Prior w

model-releasesarxiv-cs-cl
10 Jun 2026
Model Releases

Beyond Model Size: Probing the Gaps in Visual in-Context Learning by Training a Tiny Model

DGX agent

arXiv:2606.10905v1 Announce Type: new Abstract: Visual in-Context Learning (VICL) aims at making progress towards adaptive vision models, that can -- based on a few examples -- adapt to a new task at

model-releasesarxiv-cs-cv
10 Jun 2026
Model Releases

Bittensor Agent Arenas as a Trajectory Primitive: Distilling a Shopping Agent from ShoppingBench Subnet Traces

DGX agent

arXiv:2606.10064v1 Announce Type: cross Abstract: Small-model agentic post-training is bottlenecked less by the algorithm than by the trajectory substrate it consumes. Leading recipes (RLVR, group-rel

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

Boosting Graph Robustness Against Backdoor Attacks: An Over-Similarity Perspective

DGX agent

arXiv:2502.01272v3 Announce Type: replace Abstract: Graph Neural Networks (GNNs) have achieved notable success in tasks such as social and transportation networks. However, recent studies have highlig

model-releasesarxiv-cs-lg
10 Jun 2026
Model Releases

CableRobotGraphSim: A Graph Neural Network for Modeling Partially Observable Cable-Driven Robot Dynamics

DGX agent

arXiv:2602.21331v2 Announce Type: replace Abstract: General-purpose simulators have accelerated the development of robots. Traditional simulators based on first-principles, however, typically require

model-releasesarxiv-cs-ro
10 Jun 2026
Model Releases

Can Image Models Imagine Time? ImageTime: A Novel Benchmark for Probing Visual World Modeling Through Spatiotemporal Consistency

DGX agent

arXiv:2606.10620v1 Announce Type: cross Abstract: Image generation models now produce high-quality static images, yet their ability to represent how a visual world changes over time remains poorly und

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

Can Multi-Agent LLMs Identify Their Peers? Stylometric Fingerprinting in Role-Constrained Political Analysis

DGX agent

arXiv:2606.09854v1 Announce Type: cross Abstract: Multi-agent large language model (LLM) pipelines for political statement analysis are vulnerable to peer-preservation bias: models tend to protect pee

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

ChartLens: A Dual-Branch Framework for Chart Data Correction and Factual Summary Refinement

DGX agent

arXiv:2606.10640v1 Announce Type: new Abstract: In this report, we present our champion solution for the DataMFM Challenge Track 2: Chart Understanding. This track requires models to recover structure

model-releasesarxiv-cs-cv
10 Jun 2026
Model Releases

CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs

DGX agent

arXiv:2606.11063v1 Announce Type: new Abstract: AI control protocols oversee untrusted models by monitoring their actions and modifying potentially unsafe steps, often using a trusted model. This part

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

CITRAS-FM: Tiny Time Series Foundation Model for Covariate-Informed Zero-Shot Forecasting

DGX agent

arXiv:2606.10798v1 Announce Type: new Abstract: Pretrained time series foundation models (TSFMs) have enabled zero-shot forecasting on unseen target series. However, existing TSFMs often incur high co

model-releasesarxiv-cs-lg
10 Jun 2026
Model Releases

Claude Fable 5 is now available in Computer as an orchestrator model. This is Anthropic's state-of-the-art model for long, complex tasks. Av…

DGX agent

Claude Fable 5 is now available in Computer as an orchestrator model designed by Anthropic for handling long and complex tasks. This represents Anthropic's latest state-of-the-art model offering enhan

model-releasesperplexity--x
10 Jun 2026
Model Releases

Claude Fable 5 thinks document parsing is beneath it It is absolutely crushing on all reasoning-intensive/long horizon benchmarks: SWE-Bench…

DGX agent

Claude Fable 5 thinks document parsing is beneath it It is absolutely crushing on all reasoning-intensive/long horizon benchmarks: SWE-Bench Pro, FrontierCode, GDPval, Runescape, etc. But for document

model-releasesjerry-liu--x
10 Jun 2026
Model Releases

Claude Fable won’t answer basic biology questions

DGX agent

Anthropic just released Claude Fable 5, calling it the most powerful AI model it has ever made widely available and praising its skills in biology, among others. But the model won't answer basic biolo

model-releasesthe-verge-ai
10 Jun 2026
Model Releases

CleanPatrick: A Benchmark for Image Data Cleaning

DGX agent

arXiv:2505.11034v2 Announce Type: replace-cross Abstract: Robust machine learning depends on clean data, yet current image data cleaning benchmarks rely on synthetic noise or narrow human studies, lim

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

CLP: Collocation-Length Prediction for Zero-Loss Adaptive Multi-Token Inference

DGX agent

arXiv:2606.10935v1 Announce Type: cross Abstract: Large language model inference is bottlenecked by autoregressive decoding, where each token requires a full forward pass. Multi-token prediction (MTP)

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

ClusBench: The Clustering Benchmark Data Resource You've All Been Waiting For (?)

DGX agent

arXiv:2606.10673v1 Announce Type: cross Abstract: Although some very common test beds exist for assessing the performance of clustering methods, large scale benchmarking is typically limited to relati

model-releasesarxiv-cs-lg
10 Jun 2026
Model Releases

CoCoSI: Collaborative Cognitive Map Construction for Spatial Intelligence

DGX agent

arXiv:2606.10401v1 Announce Type: new Abstract: Spatial intelligence is a key frontier for multimodal large language models (MLLMs), enabling them to reason about the physical world from visual experi

model-releasesarxiv-cs-cv
10 Jun 2026
Model Releases

CodeAlchemy: Synthetic Code Rewriting at Scale

DGX agent

arXiv:2606.10087v1 Announce Type: new Abstract: Pre-training on raw code teaches syntax but provides sparse signal for diverse real-world task formats. While synthetic data has proven transformative f

model-releasesarxiv-cs-cl
10 Jun 2026
Model Releases

Cohere Transcribe, our open-source speech recognition model, is #1 on the new @huggingface Far-Field ASR benchmark.

DGX agent

Cohere has released Transcribe, an open-source speech recognition model that achieved the top ranking on Hugging Face's newly established Far-Field Automatic Speech Recognition (ASR) benchmark. The mo

model-releasescohere--x
10 Jun 2026
Model Releases

CollabSkill: Evaluating Human-Agent Collaboration On Real-World Tasks

DGX agent

arXiv:2606.09833v1 Announce Type: cross Abstract: AI agents are reshaping the workspace, leading to drastic change of how humans work. Despite the considerable potential of human-agent collaboration b

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

ComBench: A Benchmark for Rigorous Proof Reasoning and Constructive Realization in Olympiad-Level Combinatorics

DGX agent

arXiv:2606.10479v1 Announce Type: new Abstract: Combinatorics is central to Olympiad-level mathematical problem solving, requiring deep discrete reasoning, creative constructions, and rigorous structu

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data

DGX agent

arXiv:2601.18026v2 Announce Type: replace Abstract: Language identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, e

model-releasesarxiv-cs-cl
10 Jun 2026
Model Releases

Compile Once, Differentiate Everywhere: A Differentiable Meta-Circular Interpreter

DGX agent

arXiv:2606.09930v1 Announce Type: cross Abstract: The boundary between program execution and gradient-based optimization has long limited the use of code itself as a learnable scientific model. We pre

model-releasesarxiv-cs-lg
10 Jun 2026
Model Releases

Constructing coherent spatial memory in LLM agents through graph rectification

DGX agent

arXiv:2510.04195v2 Announce Type: replace Abstract: Given a map description through global traversal navigation instructions, an LLM can often infer the implicit spatial layout and answer user queries

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

Continual LLM Upcycling: A Predictor-Gated Bank-Wise Sparsity Training Recipe for Dense-to-Sparse LLMs

DGX agent

arXiv:2606.10722v1 Announce Type: new Abstract: We study dense-to-sparse continual training as a way to construct channel-sparse large language models from dense checkpoints. Starting from a Qwen2.5-8

model-releasesarxiv-cs-cl
10 Jun 2026
Model Releases

Continuous Neural Reparameterization as a Deep Geometric Prior for Robust Fixed-Chart UV Repair

DGX agent

arXiv:2606.10050v1 Announce Type: cross Abstract: Traditional UV unwrapping relies on direct optimization of geometric distortion energies and can fail through invalid initialization, local minima, or

model-releasesarxiv-cs-cv
10 Jun 2026
Model Releases

ConvMemory v2: A Recall-Preserving Top-10 Evidence Reranker for Conversational Memory Retrieval

DGX agent

arXiv:2606.10842v1 Announce Type: new Abstract: We describe ConvMemory v2, an opt-in token-evidence reranker that sits after the lightweight ConvMemory v1 reranker and reorders only v1's protected top

model-releasesarxiv-cs-cl
10 Jun 2026
Model Releases

CoTAL: Human-in-the-Loop Prompt Engineering for Generalizable Formative Assessment Scoring and Feedback

DGX agent

arXiv:2504.02323v4 Announce Type: replace Abstract: Large language models (LLMs) have created new opportunities to assist teachers and support student learning. While researchers have explored various

model-releasesarxiv-cs-cl
10 Jun 2026
Model Releases

Cybersecurity researchers complain that Claude Fable's guardrails are too strict, rejecting 'innocuous tasks' like reading blog posts or performing code reviews (Lorenzo Franceschi-Bicchierai/TechCrunch)

DGX agent

Lorenzo Franceschi-Bicchierai / TechCrunch: Cybersecurity researchers complain that Claude Fable's guardrails are too strict, rejecting “innocuous tasks” like reading blog posts or performing code rev

model-releasestechmeme
10 Jun 2026
Model Releases

Cyst-X: A Multi-Center MRI Benchmark and Federated Learning Framework for Malignancy-Risk Stratification of Pancreatic Cystic Neoplasm

DGX agent

arXiv:2507.22017v4 Announce Type: replace-cross Abstract: Pancreatic cancer is projected to be the second-deadliest cancer by 2030, making early detection critical. Intraductal papillary mucinous neop

model-releasesarxiv-cs-cv
10 Jun 2026
Model Releases

Dario Amodei says he doesn't know what role Claude played in a missile strike on an Iranian school, and its use in this instance didn't violate Anthropic's ToS (Bloomberg)

DGX agent

Bloomberg: Dario Amodei says he doesn't know what role Claude played in a missile strike on an Iranian school, and its use in this instance didn't violate Anthropic's ToS — Anthropic PBC's boss said h

model-releasestechmeme
10 Jun 2026
Model Releases

Data-Driven Dynamic Assortment in Online Platforms: Learning about Two Sides

DGX agent

arXiv:2606.11118v1 Announce Type: new Abstract: We study a dynamic assortment problem on a two-sided service platform with incomplete information and heterogeneous customers in a discrete-time setting

model-releasesarxiv-cs-lg
10 Jun 2026
Model Releases

Data-Driven Runway and Taxiway Exits Prediction of Landing Aircraft: A Case Study at Hartsfield-Jackson Atlanta International Airport

DGX agent

arXiv:2606.11017v1 Announce Type: new Abstract: Airport surface operations increasingly constrain performance at high-throughput hubs. This study examines arrival taxi-in decisions at Hartsfield-Jacks

model-releasesarxiv-cs-lg
10 Jun 2026
Model Releases

datasette-agent 0.2a0

DGX agent

Release: datasette-agent 0.2a0 Highlights from the release notes: Tools can now ask the user questions mid-execution. Tools that declare a context parameter receive a ToolContext object, and await con

model-releasessimon-willison
10 Jun 2026
Model Releases

Day 0 Anthropic Fable 5 in ParseBench: We tested the model's advancements when it comes to document understanding. The model clearly peaks w…

DGX agent

Day 0 Anthropic Fable 5 in ParseBench: We tested the model's advancements when it comes to document understanding. The model clearly peaks when it comes to adherence to the original text: 📃 Content fa

model-releasesjerry-liu--x
10 Jun 2026
Model Releases

DB-3DME: From Dataset to Benchmark for Human-aligned Automatic 3D Mesh Evaluation

DGX agent

arXiv:2606.10142v1 Announce Type: new Abstract: Recent advances in 3D generation have led to substantial improvements in realism, controllability, and efficiency, yet the evaluation of 3D assets remai

model-releasesarxiv-cs-cv
10 Jun 2026
Model Releases

Deployment-Time Memorization in Foundation-Model Agents

DGX agent

arXiv:2606.10062v1 Announce Type: new Abstract: Foundation-model agents are increasingly long-lived systems that remember users across interactions, making memorization an explicit deployment-time fun

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

Detecting Knowledge Gaps from Conversational AI Interactions Using Curriculum Prerequisite Graphs

DGX agent

arXiv:2606.10736v1 Announce Type: cross Abstract: Large online courses generate thousands of student questions directed at conversational AI teaching assistants, yet these interaction logs remain larg

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

DiffusionGemma

DGX agent

DiffusionGemma Last May Google briefly released an experimental Gemini Diffusion model. I tried the preview at the time and recorded it running at 857 tokens/second. It was an exciting model, but Goog

model-releasessimon-willison
10 Jun 2026
← Previous
1…187188189190191…472
Next →