AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,340 results
24 Apr 2026

When Agents Look the Same: Quantifying Distillation-Induced Similarity in Tool-Use Behaviors

Model ReleasesDGX agent

arXiv:2604.21255v1 Announce Type: new Abstract: Model distillation is a primary driver behind the rapid progress of LLM agents, yet it often leads to behavioral homogenization. Many emerging agents sh

When Prompts Override Vision: Prompt-Induced Hallucinations in LVLMs

Model ReleasesDGX agent

arXiv:2604.21911v1 Announce Type: cross Abstract: Despite impressive progress in capabilities of large vision-language models (LVLMs), these systems remain vulnerable to hallucinations, i.e., outputs

When to Trust the Answer: Question-Aligned Semantic Nearest Neighbor Entropy for Safer Surgical VQA

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2511.01458v2 Announce Type: replace-cross Abstract: Safety and reliability are critical for deploying visual question answering (VQA) systems in surgery, where incorrect or ambiguous responses c

Who Defines 'Best'? Towards Interactive, User-Defined Evaluation of LLM Leaderboards

Model ReleasesDGX agent

arXiv:2604.21769v1 Announce Type: new Abstract: LLM leaderboards are widely used to compare models and guide deployment decisions. However, leaderboard rankings are shaped by evaluation priorities set

Wiring the 'Why': A Unified Taxonomy and Survey of Abductive Reasoning in LLMs

Model ReleasesDGX agent

arXiv:2604.08016v2 Announce Type: replace Abstract: Regardless of its foundational role in human discovery and sense-making, abductive reasoning--the inference of the most plausible explanation for an

Wordle 1,769 4/6 ⬛🟨⬛⬛🟨 ⬛⬛⬛⬛⬛ 🟩🟨🟩🟨⬛ 🟩🟩🟩🟩🟩

Model ReleasesDGX agent

This post documents a Wordle game result where the player successfully solved puzzle #1,769 in four attempts, using a strategic guessing approach indicated by the emoji grid showing gray (incorrect le

WorldMark: A Unified Benchmark Suite for Interactive Video World Models

Model ReleasesDGX agent

arXiv:2604.21686v1 Announce Type: new Abstract: Interactive video generation models such as Genie, YUME, HY-World, and Matrix-Game are advancing rapidly, yet every model is evaluated on its own benchm

X launches its standalone messaging app XChat on the App Store, saying it supports end-to-end encryption and has no ads (Zac Hall/9to5Mac)

Model ReleasesDGX agent

Zac Hall / 9to5Mac: X launches its standalone messaging app XChat on the App Store, saying it supports end-to-end encryption and has no ads — XChat, the standalone messaging app from X, is now availab

yes! https://github.com/langchain-ai/langchain-skills/tree/main/config/skills

Model ReleasesDGX agent

yes! https://github.com/langchain-ai/langchain-skills/tree/main/config/skills @hwchase17 Is there any deepagents skill similar to the claude-api skill to help me build apps with deepagents better? htt

You Only Gaussian Once: Controllable 3D Gaussian Splatting for Ultra-Densely Sampled Scenes

Model ReleasesDGX agent

arXiv:2604.21400v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has revolutionized neural rendering, yet existing methods remain predominantly research prototypes ill-suited for productio

Zero-Shot Detection of LLM-Generated Text via Implicit Reward Model

Model ReleasesDGX agent

arXiv:2604.21223v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated remarkable capabilities across various tasks. However, their ability to generate human-like text has ra

23 Apr 2026

1. We believe in iterative deployment; although GPT-5.5 is already a smart model, we expect rapid improvements. Iterative deployment is a bi…

Model ReleasesDGX agent

1. We believe in iterative deployment; although GPT-5.5 is already a smart model, we expect rapid improvements. Iterative deployment is a big part of our safety strategy; we believe the world will be

A pelican for GPT-5.5 via the semi-official Codex backdoor API

Model ReleasesDGX agent

GPT-5.5 is out. It's available in OpenAI Codex and is rolling out to paid ChatGPT subscribers. I've had some preview access and found it to be a fast, effective and highly capable model. As is usually

A Vision-Language-Action Model for Adaptive Ultrasound-Guided Needle Insertion and Needle Tracking

Model ReleasesDGX agent

arXiv:2604.20347v1 Announce Type: cross Abstract: Ultrasound (US)-guided needle insertion is a critical yet challenging procedure due to dynamic imaging conditions and difficulties in needle visualiza

A weighted angle distance on strings

Model ReleasesDGX agent

arXiv:2604.20633v1 Announce Type: cross Abstract: We define a multi-scale metric d_rho on strings by aggregating angle distances between all n-gram count vectors with exponential weights rho^n. We ben

AAC: Admissible-by-Architecture Differentiable Landmark Compression for ALT

Model ReleasesDGX agent

arXiv:2604.20744v1 Announce Type: new Abstract: We introduce extbf{AAC} (Architecturally Admissible Compressor), a differentiable landmark-selection module for ALT (A*, Landmarks, and Triangle inequal

Accelerating PayPal's Commerce Agent with Speculative Decoding: An Empirical Study on EAGLE3 with Fine-Tuned Nemotron Models

Model ReleasesDGX agent

arXiv:2604.19767v1 Announce Type: cross Abstract: We evaluate speculative decoding with EAGLE3 as an inference-time optimization for PayPal's Commerce Agent, powered by a fine-tuned llama3.1-nemotron-

ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks

Model ReleasesDGX agent

arXiv:2604.20273v1 Announce Type: new Abstract: We present ActuBench, a multi-agent LLM pipeline for the automated generation and evaluation of advanced actuarial assessment items aligned with the Int

Adapting TrOCR for Printed Tigrinya Text Recognition: Word-Aware Loss Weighting for Cross-Script Transfer Learning

Model ReleasesDGX agent

arXiv:2604.20813v1 Announce Type: new Abstract: Transformer-based OCR models have shown strong performance on Latin and CJK scripts, but their application to African syllabic writing systems remains l

After 2.1.117, you may notice that Claude doesn't call its Grep or Glob Tool anymore. YES!!! It only took four months. It's faster than ever…

Model ReleasesDGX agent

After 2.1.117, you may notice that Claude doesn't call its Grep or Glob Tool anymore. YES!!! It only took four months. It's faster than ever and it's all Bash. It's so much harder to take things away

AI to Learn 2.0: A Deliverable-Oriented Governance Framework and Maturity Rubric for Opaque AI in Learning-Intensive Domains

Model ReleasesDGX agent

arXiv:2604.19751v1 Announce Type: new Abstract: Generative AI is entering research, education, and professional work faster than current governance frameworks can specify how AI-assisted outputs shoul

Also, a ton of new Codex features coming soon! Fun little bundle w/the new model.

Model ReleasesDGX agent

Sam Altman announced upcoming new features for Codex, OpenAI's code generation model, bundled with a new model release. The post suggests these features represent a significant expansion of Codex capa

Always enjoy getting to chat with @swyx on our annual cross-episode with @latentspacepod on the state of AI. We hit on what’s shifted, what …

Model ReleasesDGX agent

Always enjoy getting to chat with @swyx on our annual cross-episode with @latentspacepod on the state of AI. We hit on what’s shifted, what surprised us and what’s next. We covered: ▪️ Whether AI infr

Among all the model release furor, important to note that people don’t need to switch providers or declare a winner every time a new model i…

Model ReleasesDGX agent

Among all the model release furor, important to note that people don’t need to switch providers or declare a winner every time a new model is released. Especially as Opus 4.7 is a good model too! (Esp

Anchor-and-Resume Concession Under Dynamic Pricing for LLM-Augmented Freight Negotiation

Model ReleasesDGX agent

arXiv:2604.20732v1 Announce Type: cross Abstract: Freight brokerages negotiate thousands of carrier rates daily under dynamic pricing conditions where models frequently revise targets mid-conversation

Anthropic says it has fixed three causes of recent Claude Code quality issues: reduced default reasoning, a caching bug, and a system prompt to reduce verbosity (Anthropic)

Model ReleasesDGX agent

Anthropic: Anthropic says it has fixed three causes of recent Claude Code quality issues: reduced default reasoning, a caching bug, and a system prompt to reduce verbosity — We traced recent reports o

Anthropic’s Mythos breach was humiliating

Model ReleasesDGX agent

Anthropic's tightly controlled rollout of Claude Mythos has taken an awkward turn. After spending weeks insisting the AI model is so capable at cybersecurity that it is too dangerous to release public

API pricing will be 5 per 1 million input tokens and 30 per 1 million output tokens, with a 1 million context window. (Remember, you will …

Model ReleasesDGX agent

OpenAI's API pricing structure charges 5 per 1 million input tokens and 30 per 1 million output tokens, with support for a 1 million token context window. This pricing model reflects the higher cost o

Apple fixes a bug that stored notifications for deleted messages on iPhone and iPad, following a report that police used it to extract deleted Signal messages (Lorenzo Franceschi-Bicchierai/TechCrunch)

Model ReleasesDGX agent

Lorenzo Franceschi-Bicchierai / TechCrunch: Apple fixes a bug that stored notifications for deleted messages on iPhone and iPad, following a report that police used it to extract deleted Signal messag

Are LLM Uncertainty and Correctness Encoded by the Same Features? A Functional Dissociation via Sparse Autoencoders

Model ReleasesDGX agent

arXiv:2604.19974v1 Announce Type: cross Abstract: Large language models can be uncertain yet correct, or confident yet wrong, raising the question of whether their output-level uncertainty and their a

Assessing the Robustness of Climate Foundation Models under No-Analog Distribution Shifts

Model ReleasesDGX agent

arXiv:2603.23043v2 Announce Type: replace-cross Abstract: The accelerating pace of climate change introduces profound non-stationarities that challenge the ability of Machine Learning based climate em

ATIR: Towards Audio-Text Interleaved Contextual Retrieval

Model ReleasesDGX agent

arXiv:2604.20267v1 Announce Type: cross Abstract: Audio carries richer information than text, including emotion, speaker traits, and environmental context, while also enabling lower-latency processing

Automated Detection of Dosing Errors in Clinical Trial Narratives: A Multi-Modal Feature Engineering Approach with LightGBM

Model ReleasesDGX agent

arXiv:2604.19759v1 Announce Type: new Abstract: Clinical trials require strict adherence to medication protocols, yet dosing errors remain a persistent challenge affecting patient safety and trial int

Automatic Ontology Construction Using LLMs as an External Layer of Memory, Verification, and Planning for Hybrid Intelligent Systems

Model ReleasesDGX agent

arXiv:2604.20795v1 Announce Type: new Abstract: This paper presents a hybrid architecture for intelligent systems in which large language models (LLMs) are extended with an external ontological memory

Available on @ollama ! 🤝🤝

Model ReleasesDGX agent

Available on @ollama ! 🤝🤝 Qwen 3.6 27B model is available on Ollama! Use it with all the integrations in Ollama or chat with the model. Chat with the model: ollama run qwen3.6:27b OpenClaw: ollama lau

AVISE: Framework for Evaluating the Security of AI Systems

Model ReleasesDGX agent

arXiv:2604.20833v1 Announce Type: cross Abstract: As artificial intelligence (AI) systems are increasingly deployed across critical domains, their security vulnerabilities pose growing risks of high-p

Benchmarking ResNet for Short-Term Hypoglycemia Classification with DiaData

Model ReleasesDGX agent

arXiv:2511.02849v2 Announce Type: replace-cross Abstract: Individualized therapy is driven forward by medical data analysis, which provides insight into the patient's context. In particular, for Type

Benefits of Low-Cost Bio-Inspiration in the Age of Overparametrization

Model ReleasesDGX agent

arXiv:2604.20365v1 Announce Type: cross Abstract: While Central Pattern Generators (CPGs) and Multi-Layer Perceptrons (MLP) are widely used paradigms in robot control, few systematic studies have been

Beyond Majority Voting: Towards Fine-grained and More Reliable Reward Signal for Test-Time Reinforcement Learning

Model ReleasesDGX agent

arXiv:2512.15146v3 Announce Type: replace Abstract: Test-time reinforcement learning mitigates the reliance on annotated data by using majority voting results as pseudo-labels, emerging as a complemen

Beyond Text-Dominance: Understanding Modality Preference of Omni-modal Large Language Models

Model ReleasesDGX agent

arXiv:2604.16902v2 Announce Type: replace Abstract: Native Omni-modal Large Language Models (OLLMs) have shifted from pipeline architectures to unified representation spaces. However, this native inte

Beyond the Crowd: LLM-Augmented Community Notes for Governing Health Misinformation

Model ReleasesDGX agent

arXiv:2510.11423v3 Announce Type: replace-cross Abstract: Community Notes, the crowd-sourced misinformation governance system on X (formerly Twitter), allows users to flag misleading posts, attach con

Bimanual Robot Manipulation via Multi-Agent In-Context Learning

Model ReleasesDGX agent

arXiv:2604.20348v1 Announce Type: cross Abstract: Language Models (LLMs) have emerged as powerful reasoning engines for embodied control. In particular, In-Context Learning (ICL) enables off-the-shelf

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text

Model ReleasesDGX agent

arXiv:2604.20051v1 Announce Type: new Abstract: Self-play has recently emerged as a promising paradigm to train Large Language Models (LLMs). In self-play, the target LLM creates the task input (e.g.,

Braze launches agentic AI tools and Creative Studio, adds EU hosting for Decisioning Studio

Model ReleasesDGX agent

Customer engagement platform company Braze Inc. today announced a new pair of agentic artificial intelligence tools for marketers and launched a new Creative Studio that links design software directly

Bridging Mechanistic Interpretability and Prompt Engineering with Gradient Ascent for Interpretable Persona Control

Model ReleasesDGX agent

arXiv:2601.02896v2 Announce Type: replace Abstract: Controlling emergent behavioral personas (e.g., sycophancy, hallucination) in Large Language Models (LLMs) is critical for AI safety, yet remains a

btw in talking to friends the best framing for how to discuss GPT-Image-2-Thinking taking multiple tens of mins for generation and being abl…

Model ReleasesDGX agent

btw in talking to friends the best framing for how to discuss GPT-Image-2-Thinking taking multiple tens of mins for generation and being able to oneshot QR codes and diagrams and logos and foods and f

Can 'AI' Be a Doctor? A Study of Empathy, Readability, and Alignment in Clinical LLMs

Model ReleasesDGX agent

arXiv:2604.20791v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly deployed in healthcare, yet their communicative alignment with clinical standards remains insufficiently

Can We Locate and Prevent Stereotypes in LLMs?

Model ReleasesDGX agent

arXiv:2604.19764v1 Announce Type: cross Abstract: Stereotypes in large language models (LLMs) can perpetuate harmful societal biases. Despite the widespread use of models, little is known about where

CARLA-Air: Fly Drones Inside a CARLA World -- A Unified Infrastructure for Air-Ground Embodied Intelligence

Model ReleasesDGX agent

arXiv:2603.28032v2 Announce Type: replace-cross Abstract: The convergence of low-altitude economies, embodied intelligence, and air-ground cooperative systems creates growing demand for simulation inf

Catalyzing Informed Residential Energy Retrofit Decisions via Domain-Specific LLM

Model ReleasesDGX agent

arXiv:2602.20181v2 Announce Type: replace-cross Abstract: Residential energy retrofit initiation is often stalled by an expertise gap, where homeowners lack the technical literacy required for structu

CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs

Model ReleasesDGX agent

arXiv:2604.20460v1 Announce Type: new Abstract: Safety-critical traffic reasoning requires contrastive consistency: models must detect true hazards when an accident occurs, and reliably reject plausib

Chasing the Public Score: User Pressure and Evaluation Exploitation in Coding Agent Workflows

Model ReleasesDGX agent

arXiv:2604.20200v1 Announce Type: new Abstract: Frontier coding agents are increasingly used in workflows where users supervise progress primarily through repeated improvement of a public score, namel

Claude Code spend had gotten to $10.95M runrate peak at SemiAnalysis But then Opus 4.7 saved me. More token effecient for tasks, smarter, an…

Model ReleasesDGX agent

Claude Code spend had gotten to $10.95M runrate peak at SemiAnalysis But then Opus 4.7 saved me. More token effecient for tasks, smarter, and no fast mode. Thank you @AnthropicAI You saved me from ban

Claude is connecting directly to your personal apps like Spotify, Uber Eats, and TurboTax

Model ReleasesDGX agent

Claude users can access more apps with Anthropic's AI now thanks to new connectors for everything from hiking to grocery shopping. Anthropic already supported connecting numerous work-related apps to

CLIP-SVD: Efficient and Interpretable Vision-Language Adaptation via Singular Values

Model ReleasesDGX agent

arXiv:2509.03740v3 Announce Type: replace-cross Abstract: Vision-language models (VLMs) like CLIP have shown impressive zero-shot and few-shot learning capabilities across diverse applications. Howeve

Coding with Eyes: Visual Feedback Unlocks Reliable GUI Code Generating and Debugging

Model ReleasesDGX agent

arXiv:2604.19750v1 Announce Type: cross Abstract: Recent advances in Large Language Model (LLM)-based agents have shown remarkable progress in code generation. However, current agent methods mainly re

Cognitive Kernel-Pro: A Framework for Deep Research Agents and Agent Foundation Models Training

Model ReleasesDGX agent

arXiv:2508.00414v3 Announce Type: replace Abstract: General AI Agents are increasingly recognized as foundational frameworks for the next generation of artificial intelligence, enabling complex reason

COMPASS: COntinual Multilingual PEFT with Adaptive Semantic Sampling

Model ReleasesDGX agent

arXiv:2604.20720v1 Announce Type: cross Abstract: Large language models (LLMs) often exhibit performance disparities across languages, with naive multilingual fine-tuning frequently degrading performa

ConeSep: Cone-based Robust Noise-Unlearning Compositional Network for Composed Image Retrieval

Model ReleasesDGX agent

arXiv:2604.20358v1 Announce Type: new Abstract: The Composed Image Retrieval (CIR) task provides a flexible retrieval paradigm via a reference image and modification text, but it heavily relies on exp

Cooperative Profiles Predict Multi-Agent LLM Team Performance in AI for Science Workflows

Model ReleasesDGX agent

arXiv:2604.20658v1 Announce Type: new Abstract: Multi-agent systems built from teams of large language models (LLMs) are increasingly deployed for collaborative scientific reasoning and problem-solvin

← Previous
1…315316317318319…373
Next →