AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,284 results
16 Apr 2026

Evaluating Supervised Machine Learning Models: Principles, Pitfalls, and Metric Selection

Model ReleasesDGX agent

arXiv:2604.13882v1 Announce Type: new Abstract: The evaluation of supervised machine learning models is a critical stage in the development of reliable predictive systems. Despite the widespread avail

Evaluating the Evaluator: Problems with SemEval-2020 Task 1 for Lexical Semantic Change Detection

Model ReleasesDGX agent

arXiv:2604.13232v1 Announce Type: new Abstract: This discussion paper re-examines SemEval-2020 Task 1, the most influential shared benchmark for lexical semantic change detection, through a three-part

Evaluating the Formal Reasoning Capabilities of Large Language Models through Chomsky Hierarchy

Model Releases

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2604.02709v2 Announce Type: replace Abstract: The formal reasoning capabilities of LLMs are crucial for advancing automated software engineering. However, existing benchmarks for LLMs lack syste

EVE: A Domain-Specific LLM Framework for Earth Intelligence

Model ReleasesDGX agent

arXiv:2604.13071v1 Announce Type: new Abstract: We introduce Earth Virtual Expert (EVE), the first open-source, end-to-end initiative for developing and deploying domain-specialized LLMs for Earth Int

Exposia: Teaching and Assessment of Academic Writing Skills for Research Project Proposals and Peer Feedback

Model ReleasesDGX agent

arXiv:2601.06536v2 Announce Type: replace Abstract: We present Exposia, the first public dataset that connects writing and feedback in higher education, enabling research on educationally grounded com

ExpSeek: Self-Triggered Experience Seeking for Web Agents

Model ReleasesDGX agent

arXiv:2601.08605v2 Announce Type: replace Abstract: Experience intervention in web agents emerges as a promising technical paradigm, enhancing agent interaction capabilities by providing valuable insi

F-Actor: Controllable Conversational Behaviour in Full-Duplex Models

Model ReleasesDGX agent

arXiv:2601.11329v3 Announce Type: replace Abstract: Spoken conversational systems require more than accurate speech generation to have human-like conversations: to feel natural and engaging, they must

Failure Makes the Agent Stronger: Enhancing Accuracy through Structured Reflection for Reliable Tool Interactions

Model ReleasesDGX agent

arXiv:2509.18847v3 Announce Type: replace-cross Abstract: Tool-augmented large language models (LLMs) are usually trained with supervised imitation or coarse-grained reinforcement learning that optimi

Fast training of accurate physics-informed neural networks without gradient descent

Model ReleasesDGX agent

arXiv:2405.20836v3 Announce Type: replace-cross Abstract: Solving time-dependent Partial Differential Equations (PDEs) is one of the most critical problems in computational science. While Physics-Info

FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks

Model ReleasesDGX agent

arXiv:2505.19662v3 Announce Type: replace-cross Abstract: This paper introduces FieldWorkArena, a benchmark for agentic AI targeting real-world field work. With the recent increase in demand for agent

FiLM-Nav: Efficient and Generalizable Navigation via VLM Fine-tuning

Model ReleasesDGX agent

arXiv:2509.16445v2 Announce Type: replace Abstract: Enabling robotic assistants to navigate complex environments and locate objects described in free-form language is a critical capability for real-wo

Finally, we’ve been talking to lots of users about helping them make the most of Claude Code. We’ll be sharing more this week on that work, …

Model ReleasesDGX agent

Anthropic is working on improvements to help users maximize the utility of Claude Code functionality, with additional details and updates to be announced later in the same week. This announcement come

FlexGuard: Continuous Risk Scoring for Strictness-Adaptive LLM Content Moderation

Model ReleasesDGX agent

arXiv:2602.23636v3 Announce Type: replace Abstract: Ensuring the safety of LLM-generated content is essential for real-world deployment. Most existing guardrail models formulate moderation as a fixed

Flow-based Generative Modeling of Potential Outcomes and Counterfactuals

Model ReleasesDGX agent

arXiv:2505.16051v4 Announce Type: replace-cross Abstract: Predicting potential and counterfactual outcomes from observational data is central to individualized decision-making, particularly in clinica

fMRI-LM: Towards a Universal Foundation Model for Language-Aligned fMRI Understanding

Model ReleasesDGX agent

arXiv:2511.21760v3 Announce Type: replace Abstract: Recent advances in multimodal large language models (LLMs) have enabled unified reasoning across images, audio, and video, but extending such capabi

For the developers building with Claude, a direct line from the team. Follow for changelogs, API releases, community updates, and deep dives…

Model ReleasesDGX agent

This is an announcement for the official Claude Developers Twitter/X account that serves as a direct communication channel from Anthropic's team to developers using Claude's API. The account provides

For those not seeing the increase, make sure you're using Opus 4.7 with the latest Claude Code

Model ReleasesDGX agent

Claude Code users should ensure they are using Opus 4.7 with the latest updates to experience performance improvements or feature enhancements. The post suggests that users not observing expected incr

Free Geometry: Refining 3D Reconstruction from Longer Versions of Itself

Model ReleasesDGX agent

arXiv:2604.14048v1 Announce Type: new Abstract: Feed-forward 3D reconstruction models are efficient but rigid: once trained, they perform inference in a zero-shot manner and cannot adapt to the test s

From Feelings to Metrics: Understanding and Formalizing How Users Vibe-Test LLMs

Model ReleasesDGX agent

arXiv:2604.14137v1 Announce Type: new Abstract: Evaluating LLMs is challenging, as benchmark scores often fail to capture models' real-world usefulness. Instead, users often rely on ``vibe-testing'':

From Plausibility to Verifiability: Risk-Controlled Generative OCR with Vision-Language Models

Model ReleasesDGX agent

arXiv:2603.19790v3 Announce Type: replace Abstract: Modern vision-language models (VLMs) can act as generative OCR engines, yet open-ended decoding can expose rare but consequential failures. We ident

From Weights to Activations: Is Steering the Next Frontier of Adaptation?

Model ReleasesDGX agent

arXiv:2604.14090v1 Announce Type: new Abstract: Post-training adaptation of language models is commonly achieved through parameter updates or input-based methods such as fine-tuning, parameter-efficie

Functional Emotions or Situational Contexts? A Discriminating Test from the Mythos Preview System Card

Model ReleasesDGX agent

arXiv:2604.13466v1 Announce Type: cross Abstract: The Claude Mythos Preview system card deploys emotion vectors, sparse autoencoder (SAE) features, and activation verbalisers to study model internals

Gaslight, Gatekeep, V1-V3: Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation

Model ReleasesDGX agent

arXiv:2604.13803v1 Announce Type: new Abstract: Vision-language models are increasingly deployed in high-stakes settings, yet their susceptibility to sycophantic manipulation remains poorly understood

Gemini can now create personalized AI images by digging around in Google Photos

Model ReleasesDGX agent

Gemini now uses user interests and Google Photos to create personalized AI images without requiring long descriptions, allowing users to simply ask for pictures of themselves or family members. This f

Gemini can now pull from Google Photos to generate personalized images

Model ReleasesDGX agent

Google's Personal Intelligence feature, which lets Gemini pull data from apps like Google Photos to offer responses tailored to you, can now use that data and its Nano Banana 2 image model to create i

GeoBridge: A Semantic-Anchored Multi-View Foundation Model Bridging Images and Text for Geo-Localization

Model ReleasesDGX agent

arXiv:2512.02697v3 Announce Type: replace Abstract: Cross-view geo-localization infers a location by retrieving geo-tagged reference images that visually correspond to a query image. However, the trad

GLM-5.1 Tool Calling Issue Fix & Chat Template Update If you are running GLM-5.1 with vLLM/SGLang and using tool calling, please update your…

Model ReleasesDGX agent

GLM-5.1 Tool Calling Issue Fix & Chat Template Update If you are running GLM-5.1 with vLLM/SGLang and using tool calling, please update your chat template. http://huggingface.co/zai-org/GLM-5.1/blob/m

Go from blank slate to analysis with BigQuery Studio notebook gallery templates

Model ReleasesDGX agent

For many data professionals, the most daunting part of a new project isn't the complexity of the data or the sophistication of the model, it’s the 'blank slate.' Staring at an empty notebook while cre

Google’s DeepMind just released new 4B and 27B MedGemma models!

Model ReleasesDGX agent

Google DeepMind released MedGemma, a collection of medical vision-language foundation models based on Gemma 3 in 4B and 27B parameter sizes, demonstrating advanced medical understanding and reasoning

Google’s Gemini 3.1 Flash TTS model offers unparalleled control over AI voices

Model ReleasesDGX agent

Google LLC’s DeepMind artificial intelligence unit today rolled out a new text-to-speech model called Gemini 3.1 Flash TTS. Unlike its earlier, robotic predecessors, it enables users to direct the voc

Had a great time at PyCon & PyData DE. Highly recommend it. Great open-source, community-focused conference with lots of builders in the Pyt…

Model ReleasesDGX agent

Had a great time at PyCon & PyData DE. Highly recommend it. Great open-source, community-focused conference with lots of builders in the Python AI, LLM and agent space. Taking a short family break, my

Happy coding! Opus 4.7 is a significant step up. To get the most out of it, take the time to adjust your workflow to take advantage of Claud…

Model ReleasesDGX agent

Happy coding! Opus 4.7 is a significant step up. To get the most out of it, take the time to adjust your workflow to take advantage of Claude running for longer & being more agentic. It feels like a n

Here's Qwen 3.6-35B-A3B v.s. Claude Opus 4.7 for 'Generate an SVG of a flamingo riding a unicycle', in case you thought Qwen might be cheati…

Model ReleasesDGX agent

This post compares the performance of Qwen 3.6-35B-A3B and Claude Opus 4.7 models on a creative task of generating SVG code for a flamingo riding a unicycle, likely demonstrating differences in their

Hessian-Enhanced Token Attribution (HETA): Interpreting Autoregressive LLMs

Model ReleasesDGX agent

arXiv:2604.13258v1 Announce Type: new Abstract: Attribution methods seek to explain language model predictions by quantifying the contribution of input tokens to generated outputs. However, most exist

Hi! I'm here with *another launch*, it just happens to be extremely niche, nerdy, and probably only for a handful of people. In the desktop …

Model ReleasesDGX agent

Hi! I'm here with *another launch*, it just happens to be extremely niche, nerdy, and probably only for a handful of people. In the desktop app, Claude Cowork and Code now have a little Bluetooth API

Hierarchical Reinforcement Learning with Runtime Safety Shielding for Power Grid Operation

Model ReleasesDGX agent

arXiv:2604.14032v1 Announce Type: cross Abstract: Reinforcement learning has shown promise for automating power-grid operation tasks such as topology control and congestion management. However, its de

HINTBench: Horizon-agent Intrinsic Non-attack Trajectory Benchmark

Model ReleasesDGX agent

arXiv:2604.13954v1 Announce Type: new Abstract: Existing agent-safety evaluation has focused mainly on externally induced risks. Yet agents may still enter unsafe trajectories under benign conditions.

HOLY 🤯 The one and only @elder_plinius just dropped an unlocked Gemma 4 E4B, and the specs are INSANE. Look at the performance shifts: → Re…

Model ReleasesDGX agent

HOLY 🤯 The one and only @elder_plinius just dropped an unlocked Gemma 4 E4B, and the specs are INSANE. Look at the performance shifts: → Refusal rate: 98.8% down to 2.1% (!!) → Compliance: 1.2% up to

How WPP accelerates humanoid robot training 10x with G4 VMs

Model ReleasesDGX agent

Editor’s note: Today we hear from Perry Nightingale, SVP of Creative AI at WPP about the workflow that cuts training time for humanoid robots from days to minutes — plus access to the open-source code

Hybrid Retrieval for COVID-19 Literature: Comparing Rank Fusion and Projection Fusion with Diversity Reranking

Model ReleasesDGX agent

arXiv:2604.13728v1 Announce Type: cross Abstract: We present a hybrid retrieval system for COVID-19 scientific literature, evaluated on the TREC-COVID benchmark (171,332 papers, 50 expert queries). Th

I edited the intro because I realized I buried the lede originally- The 1M context window is a double-edged sword. It allows Claude to do mo…

Model ReleasesDGX agent

I edited the intro because I realized I buried the lede originally- The 1M context window is a double-edged sword. It allows Claude to do more complex tasks but it can also leads to more context pollu

I think the adaptive thinking requirement in Claude Opus 4.7 is bad in the ways that all AI effort routers are bad, but magnified by the fac…

Model ReleasesDGX agent

I think the adaptive thinking requirement in Claude Opus 4.7 is bad in the ways that all AI effort routers are bad, but magnified by the fact that there is no manual override like in ChatGPT. It regul

ID and Graph View Contrastive Learning with Multi-View Attention Fusion for Sequential Recommendation

Model ReleasesDGX agent

arXiv:2604.14114v1 Announce Type: cross Abstract: Sequential recommendation has become increasingly prominent in both academia and industry, particularly in e-commerce. The primary goal is to extract

In Claude Code the default effort is now xhigh, a new level between high and max giving finer control over the reasoning/latency tradeoff. 4…

Model ReleasesDGX agent

In Claude Code the default effort is now xhigh, a new level between high and max giving finer control over the reasoning/latency tradeoff. 4.7 thinks more, so token use runs higher than 4.6. Manage it

in the grand narrative of Meta x AI, we saw the flop (Llama 4 hurhurhur), and now we’re seeing the turn: - *more* hiring since the soup wars…

Model ReleasesDGX agent

in the grand narrative of Meta x AI, we saw the flop (Llama 4 hurhurhur), and now we’re seeing the turn: - *more* hiring since the soup wars of 2025 - Zuck literally moved in with Alexandr and Nat and

IndicDB -- Benchmarking Multilingual Text-to-SQL Capabilities in Indian Languages

Model ReleasesDGX agent

arXiv:2604.13686v1 Announce Type: new Abstract: While Large Language Models (LLMs) have significantly advanced Text-to-SQL performance, existing benchmarks predominantly focus on Western contexts and

InfiniteScienceGym: An Unbounded, Procedurally-Generated Benchmark for Scientific Analysis

Model ReleasesDGX agent

arXiv:2604.13201v1 Announce Type: new Abstract: Large language models are emerging as scientific assistants, but evaluating their ability to reason from empirical data remains challenging. Benchmarks

Introducing Claude Opus 4.7, our most capable Opus model yet. It handles long-running tasks with more rigor, follows instructions more preci…

Model ReleasesDGX agent

Introducing Claude Opus 4.7, our most capable Opus model yet. It handles long-running tasks with more rigor, follows instructions more precisely, and verifies its own outputs before reporting back. Yo

Introducing GPT-Rosalind for life sciences research

Model ReleasesDGX agent

GPT-Rosalind is an AI model developed by OpenAI specifically designed to assist with life sciences research tasks. The model is trained to help researchers with applications such as analyzing biologic

It is not well-explained, but with the adaptive switch off, I get no thinking. I can set thinking levels in Claude Code, but not in Claude C…

Model ReleasesDGX agent

It is not well-explained, but with the adaptive switch off, I get no thinking. I can set thinking levels in Claude Code, but not in Claude Cowork. AI companies keep seeming to assume that coding/techn

Joint Representation Learning and Clustering via Gradient-Based Manifold Optimization

Model ReleasesDGX agent

arXiv:2604.13484v1 Announce Type: cross Abstract: Clustering and dimensionality reduction have been crucial topics in machine learning and computer vision. Clustering high-dimensional data has been ch

KMMMU: Evaluation of Massive Multi-discipline Multimodal Understanding in Korean Language and Context

Model ReleasesDGX agent

arXiv:2604.13058v1 Announce Type: new Abstract: We introduce KMMMU, a native Korean benchmark for evaluating multimodal understanding in Korean cultural and institutional settings. KMMMU contains 3,46

KV Packet: Recomputation-Free Context-Independent KV Caching for LLMs

Model ReleasesDGX agent

arXiv:2604.13226v1 Announce Type: new Abstract: Large Language Models (LLMs) rely heavily on Key-Value (KV) caching to minimize inference latency. However, standard KV caches are context-dependent: re

L2D-Clinical: Learning to Defer for Adaptive Model Selection in Clinical Text Classification

Model ReleasesDGX agent

arXiv:2604.13285v1 Announce Type: new Abstract: Clinical text classification requires choosing between specialized fine-tuned models (BERT variants) and general-purpose large language models (LLMs), y

Language steering in latent space to mitigate unintended code-switching

Model ReleasesDGX agent

arXiv:2510.13849v3 Announce Type: replace Abstract: Multilingual Large Language Models (LLMs) often exhibit hallucinations such as unintended code-switching, reducing reliability in downstream tasks.

LaoBench: A Large-Scale Multidimensional Lao Benchmark for Large Language Models

Model ReleasesDGX agent

arXiv:2511.11334v3 Announce Type: replace Abstract: The rapid advancement of large language models (LLMs) has not been matched by their evaluation in low-resource languages, especially Southeast Asian

Learning the Cue or Learning the Word? Analyzing Generalization in Metaphor Detection for Verbs

Model ReleasesDGX agent

arXiv:2604.13713v1 Announce Type: new Abstract: Metaphor detection models achieve strong benchmark performance, yet it remains unclear whether this reflects transferable generalization or lexical memo

Leveraging LLM-GNN Integration for Open-World Question Answering over Knowledge Graphs

Model ReleasesDGX agent

arXiv:2604.13979v1 Announce Type: new Abstract: Open-world Question Answering (OW-QA) over knowledge graphs (KGs) aims to answer questions over incomplete or evolving KGs. Traditional KGQA assumes a c

LiteParse hit 4.3K+ GitHub stars in a few weeks. Today it officially joins the LlamaIndex ecosystem, with its own page at http://www.llamain…

Model ReleasesDGX agent

LiteParse hit 4.3K+ GitHub stars in a few weeks. Today it officially joins the LlamaIndex ecosystem, with its own page at http://www.llamaindex.ai/liteparse?utm_medium=socials&utm_source=twitter&utm_c

LiteParse should be the default document parser you use with any AI agent (Claude Code, Claude Cowork, OpenClaw, Codex, and more) The core i…

Model ReleasesDGX agent

LiteParse should be the default document parser you use with any AI agent (Claude Code, Claude Cowork, OpenClaw, Codex, and more) The core is extremely fast text and accurate parsing from any document

← Previous
1…340341342343344…372
Next →