AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,084 results
10 Apr 2026

Tight Convergence Rates for Online Distributed Linear Estimation with Adversarial Measurements

Model ReleasesDGX agent

arXiv:2604.06282v1 Announce Type: cross Abstract: We study mean estimation of a random vector X in a distributed parameter-server-worker setup. Worker i observes samples of a_i^op X, where $a_

Tool Retrieval Bridge: Aligning Vague Instructions with Retriever Preferences via Bridge Model

Model ReleasesDGX agent

arXiv:2604.07816v1 Announce Type: new Abstract: Tool learning has emerged as a promising paradigm for large language models (LLMs) to address real-world challenges. Due to the extensive and irregularl

Toward a universal foundation model for graph-structured data

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.06391v1 Announce Type: cross Abstract: Graphs are a central representation in biomedical research, capturing molecular interaction networks, gene regulatory circuits, cell--cell communicati

Toward Memory-Aided World Models: Benchmarking via Spatial Consistency

Model ReleasesDGX agent

arXiv:2505.22976v2 Announce Type: replace-cross Abstract: The ability to simulate the world in a spatially consistent manner is a crucial requirements for effective world models. Such a model enables

Towards Accurate and Calibrated Classification: Regularizing Cross-Entropy From A Generative Perspective

Model ReleasesDGX agent

arXiv:2604.06689v1 Announce Type: new Abstract: Accurate classification requires not only high predictive accuracy but also well-calibrated confidence estimates. Yet, modern deep neural networks (DNNs

Towards Effective Long Video Understanding of Multimodal Large Language Models via One-shot Clip Retrieval

Model ReleasesDGX agent

arXiv:2512.08410v2 Announce Type: replace Abstract: Due to excessive memory overhead, most Multimodal Large Language Models (MLLMs) can only process videos of limited frames. In this paper, we propose

Towards Real-world Human Behavior Simulation: Benchmarking Large Language Models on Long-horizon, Cross-scenario, Heterogeneous Behavior Traces

Model ReleasesDGX agent

arXiv:2604.08362v1 Announce Type: new Abstract: The emergence of Large Language Models (LLMs) has illuminated the potential for a general-purpose user simulator. However, existing benchmarks remain co

ToxReason: A Benchmark for Mechanistic Chemical Toxicity Reasoning via Adverse Outcome Pathway

Model ReleasesDGX agent

arXiv:2604.06264v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) have enabled molecular reasoning for property prediction. However, toxicity arises from complex biolog

TR-EduVSum: A Turkish-Focused Dataset and Consensus Framework for Educational Video Summarization

Model ReleasesDGX agent

arXiv:2604.07553v1 Announce Type: new Abstract: This study presents a framework for generating the gold-standard summary fully automatically and reproducibly based on multiple human summaries of Turki

TraceSafe: A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories

Model ReleasesDGX agent

arXiv:2604.07223v1 Announce Type: cross Abstract: As large language models (LLMs) evolve from static chatbots into autonomous agents, the primary vulnerability surface shifts from final outputs to int

Training Data Size Sensitivity in Unsupervised Rhyme Recognition

Model ReleasesDGX agent

arXiv:2604.08156v1 Announce Type: new Abstract: Rhyme is deceptively intuitive: what is or is not a rhyme is constructed historically, scholars struggle with rhyme classification, and people disagree

TSUBASA: Improving Long-Horizon Personalization via Evolving Memory and Self-Learning with Context Distillation

Model ReleasesDGX agent

arXiv:2604.07894v1 Announce Type: new Abstract: Personalized large language models (PLLMs) have garnered significant attention for their ability to align outputs with individual's needs and preference

Ultraplan uses roughly the same number of tokens (and subscription rate limits) as plan mode. See the docs for more: http://docs.claude.com/…

Model ReleasesDGX agent

Ultraplan, a planning feature in Claude, consumes approximately the same number of tokens and counts against subscription rate limits similarly to standard plan mode. Users should be aware that using

Unifying Speech Editing Detection and Content Localization via Prior-Enhanced Audio LLMs

Model ReleasesDGX agent

arXiv:2601.21463v2 Announce Type: replace-cross Abstract: Existing speech editing detection (SED) datasets are predominantly constructed using manual splicing or limited editing operations, resulting

UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding

Model ReleasesDGX agent

arXiv:2604.08522v1 Announce Type: new Abstract: Video temporal grounding (VTG) is typically tackled with dataset-specific models that transfer poorly across domains and query styles. Recent efforts to

Validated Intent Compilation for Constrained Routing in LEO Mega-Constellations

Model ReleasesDGX agent

arXiv:2604.07264v1 Announce Type: cross Abstract: Operating LEO mega-constellations requires translating high-level operator intents ('reroute financial traffic away from polar links under 80 ms') int

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents

Model ReleasesDGX agent

arXiv:2603.15118v2 Announce Type: replace Abstract: We introduce VAREX (VARied-schema EXtraction), a benchmark for evaluating multimodal foundation models on structured data extraction from government

Variational Feature Compression for Model-Specific Representations

Model ReleasesDGX agent

arXiv:2604.06644v1 Announce Type: cross Abstract: As deep learning inference is increasingly deployed in shared and cloud-based settings, a growing concern is input repurposing, in which data submitte

VenusBench-Mobile: A Challenging and User-Centric Benchmark for Mobile GUI Agents with Capability Diagnostics

Model ReleasesDGX agent

arXiv:2604.06182v1 Announce Type: cross Abstract: Existing online benchmarks for mobile GUI agents remain largely app-centric and task-homogeneous, failing to reflect the diversity and instability of

Verify Before You Commit: Towards Faithful Reasoning in LLM Agents via Self-Auditing

Model ReleasesDGX agent

arXiv:2604.08401v1 Announce Type: cross Abstract: In large language model (LLM) agents, reasoning trajectories are treated as reliable internal beliefs for guiding actions and updating memory. However

VertAX: a differentiable vertex model for learning epithelial tissue mechanics

Model ReleasesDGX agent

arXiv:2604.06896v1 Announce Type: new Abstract: Epithelial tissues dynamically reshape through local mechanical interactions among cells, a process well captured by vertex models. Yet their many tunab

Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs

Model ReleasesDGX agent

arXiv:2509.08016v2 Announce Type: replace Abstract: Video Large Language Models (VideoLLMs) face a critical bottleneck: increasing the number of input frames to capture fine-grained temporal detail le

VisCoder2: Building Multi-Language Visualization Coding Agents

Model ReleasesDGX agent

arXiv:2510.23642v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have recently enabled coding agents capable of generating, executing, and revising visualization code. However, e

Vision-Language Navigation for Aerial Robots: Towards the Era of Large Language Models

Model ReleasesDGX agent

arXiv:2604.07705v1 Announce Type: new Abstract: Aerial vision-and-language navigation (Aerial VLN) aims to enable unmanned aerial vehicles (UAVs) to interpret natural language instructions and autonom

Visual prompting reimagined: The power of the Activation Prompts

Model ReleasesDGX agent

arXiv:2604.06440v1 Announce Type: cross Abstract: Visual prompting (VP) has emerged as a popular method to repurpose pretrained vision models for adaptation to downstream tasks. Unlike conventional mo

Visually-grounded Humanoid Agents

Model ReleasesDGX agent

arXiv:2604.08509v1 Announce Type: new Abstract: Digital human generation has been studied for decades and supports a wide range of real-world applications. However, most existing systems are passively

VSAS-BENCH: Real-Time Evaluation of Visual Streaming Assistant Models

Model ReleasesDGX agent

arXiv:2604.07634v1 Announce Type: new Abstract: Streaming vision-language models (VLMs) continuously generate responses given an instruction prompt and an online stream of input frames. This is a core

WASD: Locating Critical Neurons as Sufficient Conditions for Explaining and Controlling LLM Behavior

Model ReleasesDGX agent

arXiv:2603.18474v2 Announce Type: replace Abstract: Precise behavioral control of large language models (LLMs) is critical for complex applications. However, existing methods often incur high training

We love seeing what you’ve built with Gemma 4, the open model family that we released last week. Here are a few fun examples, described by t…

Model ReleasesDGX agent

Google AI's Gemma 4, released on April 2, 2026, is Google DeepMind's most capable open model family to date, purpose-built for advanced reasoning and agentic workflows . The models are available ...

We pit LlamaParse against frontier models (Opus 4.6, Gemini 3.1 Pro, GPT-5.4) in a live OCR arena. ICYMI: the full workshop is on Youtube! F…

Model ReleasesDGX agent

We pit LlamaParse against frontier models (Opus 4.6, Gemini 3.1 Pro, GPT-5.4) in a live OCR arena. ICYMI: the full workshop is on Youtube! Frontier VLMs are getting quite good at visual understanding,

We worked with @RWSGroup to fine-tune our Command translation model, which improved language and cultural expertise. Now, that model is the …

Model ReleasesDGX agent

We worked with @RWSGroup to fine-tune our Command translation model, which improved language and cultural expertise. Now, that model is the “brain” that powers RWS’ Language Weaver AI translation solu

Weakly Supervised Distillation of Hallucination Signals into Transformer Representations

Model ReleasesDGX agent

arXiv:2604.06277v1 Announce Type: new Abstract: Existing hallucination detection methods for large language models (LLMs) rely on external verification at inference time, requiring gold answers, retri

Weighted Bayesian Conformal Prediction

Model ReleasesDGX agent

arXiv:2604.06464v1 Announce Type: new Abstract: Conformal prediction provides distribution-free prediction intervals with finite-sample coverage guarantees, and recent work by Snell & Griffiths refra

what he said 🗣️ the very best agents today obsessively tailor the harness layer around the model I’m looking at you “5 things I learned fro…

Model ReleasesDGX agent

what he said 🗣️ the very best agents today obsessively tailor the harness layer around the model I’m looking at you “5 things I learned from the Claude code leak” bros 👀 orchestration patterns, tool d

What's Missing in Screen-to-Action? Towards a UI-in-the-Loop Paradigm for Multimodal GUI Reasoning

Model ReleasesDGX agent

arXiv:2604.06995v1 Announce Type: new Abstract: Existing Graphical User Interface (GUI) reasoning tasks remain challenging, particularly in UI understanding. Current methods typically rely on direct s

What’s new with Google Cloud

Model ReleasesDGX agent

Want to know the latest from Google Cloud? Find it here in one handy location. Check back regularly for our newest updates, announcements, resources, events, learning opportunities, and more. Tip: Not

When Personalization Tricks Detectors: The Feature-Inversion Trap in Machine-Generated Text Detection

Model ReleasesDGX agent

arXiv:2510.12476v2 Announce Type: replace Abstract: Large language models (LLMs) have grown more powerful in language generation, producing fluent text and even imitating personal style. Yet, this abi

When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't

Model ReleasesDGX agent

arXiv:2604.06422v1 Announce Type: cross Abstract: Understanding when Vision-Language Models (VLMs) will behave unexpectedly, whether models can reliably predict their own behavior, and if models adher

When to Trust Tools? Adaptive Tool Trust Calibration For Tool-Integrated Math Reasoning

Model ReleasesDGX agent

arXiv:2604.08281v1 Announce Type: new Abstract: Large reasoning models (LRMs) have achieved strong performance enhancement through scaling test time computation, but due to the inherent limitations of

Which Way Does Time Flow? A Psychophysics-Grounded Evaluation for Vision-Language Models

Model ReleasesDGX agent

arXiv:2510.26241v5 Announce Type: replace-cross Abstract: Modern vision-language models (VLMs) excel at many multimodal tasks, yet their grasp of temporal information in video remains weak and has not

Why Are We Lonely? Leveraging LLMs to Measure and Understand Loneliness in Caregivers and Non-caregivers

Model ReleasesDGX agent

arXiv:2604.07834v1 Announce Type: new Abstract: This paper presents an LLM-driven approach for constructing diverse social media datasets to measure and compare loneliness in the caregiver and non-car

Wordle 1,756 5/6 ⬛⬛🟨🟨⬛ ⬛⬛⬛⬛🟨 ⬛🟨🟨🟨⬛ 🟨🟨🟨🟩⬛ 🟩🟩🟩🟩🟩

Model ReleasesDGX agent

On April 10, 2026, Anthropic posted their Wordle 1,756 result on X (Twitter), completing the puzzle in 5 out of 6 guesses. The answer to that day's puzzle was **CAROM** — a noun and verb meaning, ...

Wrapped up the 3-day AI Engineer conference in London. Key takeaways: Planning & verification > implementation. Humans define what to build …

Model ReleasesDGX agent

Wrapped up the 3-day AI Engineer conference in London. Key takeaways: Planning & verification > implementation. Humans define what to build and validate it works. AI handles the coding. The harness ma

Yet another illustration of why LLMs aren’t even close to being AGI.

Model ReleasesDGX agent

Yet another illustration of why LLMs aren’t even close to being AGI. The world’s best LLMs are still terrible at poker. We put each model into a 200bb heads-up NLHE match against GTO Wizard AI. The be

You Point, I Learn: Online Adaptation of Interactive Segmentation Models for Handling Distribution Shifts in Medical Imaging

Model ReleasesDGX agent

arXiv:2503.06717v3 Announce Type: replace Abstract: Interactive segmentation uses real-time user inputs, such as mouse clicks, to iteratively refine model predictions. Although not originally designed

9 Apr 2026

A bit over a decade ago, we got fuzzers. A fuzzer is an automated vulnerability-finder that repeatedly runs a target program with semi-rando…

Model ReleasesDGX agent

A bit over a decade ago, we got fuzzers. A fuzzer is an automated vulnerability-finder that repeatedly runs a target program with semi-random inputs. One particular fuzzer, American Fuzzy Lop, was not

A developer’s guide to architecting reliable GPU infrastructure at scale

Model ReleasesDGX agent

Editor’s note: This blog post outlines Google Cloud’s GPU AI/ML infrastructure reliability strategy, and will be updated with links to new community articles as they appear. As we enter the era of mul

After playing with it a bit, Meta's Muse Spark Thinking is fine so far, but really doesn't match the current Big Three models. It also is a …

Model ReleasesDGX agent

After playing with it a bit, Meta's Muse Spark Thinking is fine so far, but really doesn't match the current Big Three models. It also is a bit... weird. Like some strange language & tone, a little lo

AI on the couch: Anthropic gives Claude 20 hours of psychiatry

Model ReleasesDGX agent

As part of the evaluation of its Claude Mythos model, Anthropic engaged a clinical psychiatrist for approximately 20 hours of evaluation sessions, describing Mythos as 'the most psychologically se...

Apiiro launches command-line interface to bring AI-native security into software development workflows

Model ReleasesDGX agent

Application security posture management company Apiiro Ltd. today announced the launch of a new command-line interface designed to bring application security directly into artificial intelligence-driv

Appknox launches KnoxIQ to prioritize real-world exploitability in AI-driven application security

Model ReleasesDGX agent

Mobile app security solutions provider Appknox today announced the launch of KnoxIQ, an artificial intelligence-native vulnerability assessment capability that introduces a new prioritization and reme

Blaize launches AI Services platform to move enterprise AI from pilot to production

Model ReleasesDGX agent

Artificial intelligence computing company Blaize Holdings Inc. today announced the launch of Blaize AI Services, a new platform designed to help AI infrastructure providers and enterprises deploy prod

Claude Agent SDK tracing in LangSmith just got an upgrade. Now you can trace: → Subagents → Child runs inside MCP tools → Cost tracking + mo…

Model ReleasesDGX agent

Claude Agent SDK tracing in LangSmith just got an upgrade. Now you can trace: → Subagents → Child runs inside MCP tools → Cost tracking + more Update to the latest Python SDK to try it out. Docs: http

Claude Cowork, now generally available!

Model ReleasesDGX agent

Claude Cowork, now generally available! We released Claude Cowork as a research preview 12 weeks ago. Since then, millions of people have made it part of how they work - and hundreds of thousands more

Claude Mythos and misguided open-weight fearmongering

Model ReleasesDGX agent

The *Interconnects.ai* article argues that fears around releasing an open-weight version of Claude Mythos are overstated, noting that closed frontier models still lead open-weight ones in robust, ...

Day 1 of @aiDotEngineer in London – what a ride! 🤖 Ash Prabaker & Andrew Wilson from Anthropic on 'How to Build Agents That Run for Hours (…

Model ReleasesDGX agent

Day 1 of @aiDotEngineer in London – what a ride! 🤖 Ash Prabaker & Andrew Wilson from Anthropic on 'How to Build Agents That Run for Hours (Without Losing the Plot)' – spoiler: stop letting your agents

Deep Agents or @LangChain? Just use the right tool for the right job

Model ReleasesDGX agent

Deep Agents or @LangChain? Just use the right tool for the right job great q! deepagents has more 'batteries included', which langchain v1 is a very minimalistic agent harness if you are doing more co

Deploy agents with your choice of model, sandbox, and integrations

Model ReleasesDGX agent

The specific X (Twitter) post at the provided URL could not be retrieved — the content is not publicly indexed in search results, and direct access to X posts requires authentication or is otherwis...

Enjoy faster @-mentions in Claude Code! This went out in v2.1.85.

Model ReleasesDGX agent

Claude Code v2.1.85 introduced performance improvements to the `@`-mention feature, with a key optimization being that raw string content for `@`-mentioned files is no longer JSON-escaped, reducing...

every engineer at anthropic has been using mythos for ~1.5 months. meanwhile, their uptime is horrendous, claude code still has rendering bu…

Model ReleasesDGX agent

every engineer at anthropic has been using mythos for ~1.5 months. meanwhile, their uptime is horrendous, claude code still has rendering bugs, etc. one could conclude that it won't be the end of soft

← Previous
1…364365366367368369
Next →