AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,619Total entries
1Added by human
84,618Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,595 results
23 May 2026

The Illusion of Reasoning: Exposing Evasive Data Contamination in LLMs via Zero-CoT Truncation

Model ReleasesDGX agent

arXiv:2605.21856v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated impressive reasoning abilities across a wide range of tasks, but data contamination undermines the object

The Secretary Problem with a Stochastic Precursor

Model ReleasesDGX agent

arXiv:2605.22653v1 Announce Type: cross Abstract: In learning-augmented online algorithms, predictions are usually valued for what they say: a value estimate, a solution, or an algorithmic recommendat

The Volterra signature

Model ReleasesDGX agent

arXiv:2603.04525v2 Announce Type: replace-cross Abstract: Modern approaches for learning from non-Markovian time series, such as recurrent neural networks, neural controlled differential equations or


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

this has been my experience as well. there definitely were improvements, specifically wrt shell based computer use, but also regressions, es…

Model ReleasesDGX agent

this has been my experience as well. there definitely were improvements, specifically wrt shell based computer use, but also regressions, especially in the last 3 version bumps of flicker and gerperte

this is even easier

Model ReleasesDGX agent

this is even easier oh even better, since you can test the runtime without an LLM, you can post this into Claude: “look up http://activegraph.ai, install and run a small experiment I would like levera

Towards Speed-of-Light Text Generation with Nemotron-Labs Diffusion Language Models

Model ReleasesDGX agent

Nemotron-Labs Diffusion Language Models represent NVIDIA's approach to achieving faster text generation through diffusion-based architectures, potentially offering significant speed improvements over

Truncated Neural Likelihood Estimation for Simulation-Based Inference in State-Space Models

Model ReleasesDGX agent

arXiv:2605.21805v1 Announce Type: cross Abstract: State-space models (SSMs) are powerful probabilistic tools for modeling time-varying systems with latent dynamics. Inference in SSMs involves the esti

UNAD+: An Explainable Hybrid Framework for Unknown Network Attack Detection

Model ReleasesDGX agent

arXiv:2605.22621v1 Announce Type: cross Abstract: The detection of previously unseen network attacks remains a major challenge for intrusion detection systems. Although supervised learning methods oft

VeriScale: Adversarial Test-Suite Scaling for Verifiable Code Generation

Model ReleasesDGX agent

arXiv:2605.22368v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly deployed for software engineering, constructing high-quality benchmarks is crucial for evaluating not j

When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning

Model ReleasesDGX agent

arXiv:2605.21606v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) trains a student on its own rollouts using a privileged teacher, but its standard objective weights all generated tok

Why SGD is not Brownian Motion: A New Perspective on Stochastic Dynamics

Model ReleasesDGX agent

arXiv:2605.22644v1 Announce Type: new Abstract: Stochastic Gradient Descent (SGD) is commonly modeled as a Langevin process, assuming that minibatch noise acts as Brownian motion. However, this approx

Wordle 1,798 5/6 ⬛⬛🟨⬛⬛ 🟨⬛⬛⬛🟨 ⬛🟩⬛🟩🟩 ⬛🟩🟩🟩🟩 🟩🟩🟩🟩🟩

Model ReleasesDGX agent

This is a Wordle game result post from Anthropic's X (Twitter) account showing the solution found in 5 of 6 attempts, with the color-coded emoji grid indicating which letters were correct, misplaced,

22 May 2026

3D LULC classification using multispectral LiDAR and deep learning: current and prospective schemes

Model ReleasesDGX agent

arXiv:2605.22328v1 Announce Type: new Abstract: Land Use Land Cover (LULC) classification is essential for national 3D mapping, geospatial analysis, and sustainable planning. Multispectral (MS) LiDAR

A Comparative Study of Language Models for Khmer Retrieval-Augmented Question Answering

Model ReleasesDGX agent

arXiv:2605.22099v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) has emerged as a promising paradigm for grounding large language model (LLM) outputs in retrieved evidence, thereby

A Task-Agnostic Algebraic Integrity Metric for Event-Camera Streams Toward SOTIF-Compliant Perception using Pearson Correlation Coefficient

Model ReleasesDGX agent

arXiv:2605.21500v1 Announce Type: cross Abstract: Event cameras have emerged as a high-bandwidth, low-latency sensing modality for safety-critical perception in automated driving systems (ADS), offeri

AesFormer: Transform Everyday Photos into Beautiful Memories

Model ReleasesDGX agent

arXiv:2605.22126v1 Announce Type: new Abstract: In everyday photography, aesthetically appealing moments are often captured with structural flaws (e.g., composition, camera viewpoint, or pose) that ex

AgentCo-op: Retrieval-Based Synthesis of Interoperable Multi-Agent Workflows

Model ReleasesDGX agent

arXiv:2605.20425v1 Announce Type: new Abstract: Designing multi-agent workflows is especially difficult in open-ended scientific settings where tasks lack curated training sets, reliable scalar evalua

AgroTools: A Benchmark for Tool-Augmented Multimodal Agents in Agriculture

Model ReleasesDGX agent

arXiv:2605.22366v1 Announce Type: new Abstract: Agricultural decision-making increasingly requires multimodal systems that can transform visual observations into reliable, executable actions. However,

AgroVG: A Large-Scale Multi-Source Benchmark for Agricultural Visual Grounding

Model ReleasesDGX agent

arXiv:2605.22034v1 Announce Type: new Abstract: Visual grounding, the task of localizing objects described by natural-language expressions, is a foundational capability for agricultural AI systems, en

AlignPose: Generalizable 6D Pose Estimation via Multi-view Feature-metric Alignment

Model ReleasesDGX agent

arXiv:2512.20538v2 Announce Type: replace Abstract: Single-view RGB model-based object pose estimation methods achieve strong generalization but are fundamentally limited by depth ambiguity, clutter,

AMEL: Accumulated Message Effects on LLM Judgments

Model ReleasesDGX agent

arXiv:2605.22714v1 Announce Type: cross Abstract: Large language models are routinely used as automated evaluators: to review code, moderate content, or score outputs, often with many items passing th

Anthropic says Claude Mythos Preview has been used to find more than 10,000 high- or critical-severity vulnerabilities since the launch of Project Glasswing (Anthropic)

Model ReleasesDGX agent

Anthropic: Anthropic says Claude Mythos Preview has been used to find more than 10,000 high- or critical-severity vulnerabilities since the launch of Project Glasswing — Last month, we launched Projec

ArabDiscrim: A Decade-Long Arabic Facebook Corpus on Racism and Discrimination

Model ReleasesDGX agent

arXiv:2605.22081v1 Announce Type: new Abstract: We present ArabDiscrim, a decade-long lexical resource and corpus of 293K public Arabic Facebook posts (2014--2024) discussing racism and discrimination

Audience Engagement with Arabic Women's Social Empowerment and Wellbeing: A Decadal Corpus

Model ReleasesDGX agent

arXiv:2605.22204v1 Announce Type: new Abstract: This paper presents the Arabic Women and Society Corpus, a ten year collection of 252,487 public Arabic Facebook posts related to women's empowerment an

BEiTScore: Reference-free Image Captioning Evaluation with an Efficient Cross-Encoder Model

Model ReleasesDGX agent

arXiv:2605.21728v1 Announce Type: cross Abstract: Image captioning evaluation remains a significant challenge, as vision-language models evolve toward more challenging capabilities such as generating

Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Models

Model ReleasesDGX agent

arXiv:2605.22732v1 Announce Type: cross Abstract: We investigate whether acoustic emotion recognition models can serve as proxies for the Pathos dimension in political speech analysis, as operationali

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI

Model ReleasesDGX agent

arXiv:2603.14987v2 Announce Type: replace Abstract: Agentic AI systems increasingly act through tool-augmented, multi-step workflows whose failures (unsafe tool use, unauthorised actions, social harm)

Beyond Chamfer Distance: Granular Order-aware Evaluation Metric For Online Mapping

Model ReleasesDGX agent

arXiv:2605.22578v1 Announce Type: new Abstract: Online map estimation is a crucial component of autonomous driving systems that reduces the reliance on costly high-definition maps. State-of-the-art (S

Beyond Euclidean Proximity: Repairing Latent World Models with Horizon-Matched Trajectory Reachability Metrics

Model ReleasesDGX agent

arXiv:2605.22164v1 Announce Type: cross Abstract: Latent world models can contain the state needed for control, yet their terminal-cost interface can expose the planner to the wrong decision-relevant

Beyond Temperature: Hyperfitting as a Late-Stage Geometric Expansion

Model ReleasesDGX agent

arXiv:2605.22579v1 Announce Type: new Abstract: Recent work has identified a counterintuitive phenomenon termed 'Hyperfitting', where fine-tuning Large Language Models (LLMs) to near-zero training los

Big fan of teaching more people the basics of using Claude Code in an accessible way. So much of the world has not yet used agents. There's …

Model ReleasesDGX agent

Big fan of teaching more people the basics of using Claude Code in an accessible way. So much of the world has not yet used agents. There's a lot of opportunity to level the playing field and expand a

Blind Spots in the Guard: How Domain-Camouflaged Injection Attacks Evade Detection in Multi-Agent LLM Systems

Model ReleasesDGX agent

arXiv:2605.22001v1 Announce Type: cross Abstract: Injection detectors deployed to protect LLM agents are calibrated on static, template-based payloads that announce themselves as override directives.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety

Model ReleasesDGX agent

arXiv:2605.22643v1 Announce Type: new Abstract: Background. Traditional safety benchmarks for language models evaluate generated text: whether a model outputs toxic language, reproduces bias, or follo

Catch up on the Dialogues stage at Google I/O 2026.

Model ReleasesDGX agent

The Dialogues stage at Google I/O 2026 brought together Google leaders, scientific minds and creative visionaries to discuss technological breakthroughs. Featured discussions included AI agents and pr

Check Your LLM's Secret Dictionary! Five Lines of Code Reveal What Your LLM Learned (Including What It Shouldn't Have)

Model ReleasesDGX agent

arXiv:2605.22005v1 Announce Type: cross Abstract: We show that singular value decomposition of the lm_head} weight matrix of a transformer-based large language model -- requiring only five lines of Py

ChronoMedKG: A Temporally-Grounded Biomedical Knowledge Graph and Benchmark for Clinical Reasoning

Model ReleasesDGX agent

arXiv:2605.22734v1 Announce Type: new Abstract: Biomedical knowledge graphs (KGs) treat disease associations as static facts, but temporal information is crucial for clinical reasoning, e.g., a sympto

Closed and gated surveillance economy SaaS systems and models will get replaced by open source in two tiers: - The models themselves - The e…

Model ReleasesDGX agent

Closed and gated surveillance economy SaaS systems and models will get replaced by open source in two tiers: - The models themselves - The everything-ultra-app harness If an American open source champ

COCOTree: A Dataset and Benchmark for Open Tree-Structured Visual Decomposition

Model ReleasesDGX agent

arXiv:2605.22068v1 Announce Type: new Abstract: We formalize and enable the task of open tree decomposition, which segments an image into hierarchical trees of visual components with unconstrained gra

Code Researcher: Deep Research Agent for Large Systems Code and Commit History

Model ReleasesDGX agent

arXiv:2506.11060v2 Announce Type: replace-cross Abstract: Large Language Model (LLM)-based coding agents have shown promising results on coding benchmarks, but their effectiveness on systems code rema

Cohesion-6K: An Arabic Dataset for Analyzing Social Cohesion and Conflict in Online Discourse

Model ReleasesDGX agent

arXiv:2605.22447v1 Announce Type: new Abstract: The study of online discourse has become central to understanding societal polarization. While much research has focused on detecting overt toxicity, th

Comparing LLM and Fine-Tuned Model Performance on NVDRS Circumstance Extraction with Varying Prompt Complexity

Model ReleasesDGX agent

arXiv:2605.21845v1 Announce Type: new Abstract: Suicide is a leading cause of death in the United States, and understanding the circumstances that precede it requires extracting structured information

CoRMA: Contrastive RMA for Contact-Rich Meta-Adaptation

Model ReleasesDGX agent

arXiv:2605.22082v1 Announce Type: new Abstract: We present CoRMA(Contrastive Robotic Motor Adaptation), a context-based meta-adaptation framework that modifies RMA for force-dominant assembly. CoRMA r

Cross-Lingual Consensus: Aligning Multilingual Cultural Knowledge via Multilingual Self-Consistency

Model ReleasesDGX agent

arXiv:2605.22137v1 Announce Type: new Abstract: Although Large Language Models (LLMs) demonstrate strong capabilities across various tasks, they exhibit significant performance discrepancies across la

CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2605.21854v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have rapidly converged on a small set of architectural patterns: discrete-token autoregression (e.g. OpenVLA) and co

CTFExplorer: Evaluating LLM Offensive Agents Through Multi-Target Web CTF Benchmarking

Model ReleasesDGX agent

arXiv:2602.08023v3 Announce Type: replace-cross Abstract: Existing benchmarks for LLM-based offensive security agents use isolated, single-target setups with a known vulnerable service and fixed objec

Cursor Composer 2.5's is 3–18x cheaper than Opus 4.7 in Claude Code (medium reasoning), and 5–32x cheaper than GPT-5.5 in Codex (medium) bas…

Model ReleasesDGX agent

Cursor Composer 2.5's is 3–18x cheaper than Opus 4.7 in Claude Code (medium reasoning), and 5–32x cheaper than GPT-5.5 in Codex (medium) based on API pricing This low Cost per Task isn't just driven b

Declarative Data Services: Structured Agentic Discovery for Composing Data Systems

Model ReleasesDGX agent

arXiv:2605.20690v1 Announce Type: new Abstract: Agentic discovery has shown that LLM-driven search can find novel algorithms, designs, and code under benchmark conditions. Translating the paradigm to

DeepSeek v4: the most expected open-source model ever released, and the quietest landing

Model ReleasesDGX agent

After 15 months of incremental updates, leaks, and rumored leaks, DeepSeek released version 4. It arrived without the fanfare R1 and R1-preview commanded in early 2025. That quiet reception is the mos

DeepSeek’s New AI Is A Game Changer

Model ReleasesDGX agent

DeepSeek released two advanced AI models—R1 and V3—with R1 specializing in complex reasoning and V3 designed for large-scale language processing. Notably, DeepSeek developed these high-performing mode

DeepWeb-Bench: A Deep Research Benchmark Demanding Massive Cross-Source Evidence and Long-Horizon Derivation

Model ReleasesDGX agent

arXiv:2605.21482v1 Announce Type: new Abstract: Deep research, in which an agent searches the open web, collects evidence, and derives an answer through extended reasoning, is a prominent use case for

DeFacto: Counterfactual Thinking with Images for Enforcing Evidence-Grounded and Faithful Reasoning

Model ReleasesDGX agent

arXiv:2509.20912v4 Announce Type: replace Abstract: Recent advances in multimodal language models (MLLMs) have made thinking with images a dominant paradigm for multimodal reasoning. However, existing

Dissecting Embodied Abilities in Multimodal Language Models through Skill-level Evaluation and Diagnosis

Model ReleasesDGX agent

arXiv:2510.08759v2 Announce Type: replace Abstract: Understanding the capability bottlenecks of embodied multimodal large language models (MLLMs) is crucial for improving embodied agents. However, exi

Diverse Yet Consistent: Context-Guided Diffusion with Energy-Based Joint Refinement for Multi-Agent Motion Prediction

Model ReleasesDGX agent

arXiv:2605.22017v1 Announce Type: new Abstract: Deepgenerative models havebecomeapromisingapproach for human motion prediction due to their ability to capture multimodal distributions and represent di

Do Vision Models Encode Object-Level Semantic Relatedness? A Cognitive Psychology-Inspired Benchmark

Model ReleasesDGX agent

arXiv:1709.03806v2 Announce Type: replace Abstract: Modern vision models have achieved strong object-recognition performance, yet it remains unclear whether their representations encode object-level s

Does Slightly Mean Somewhat? Measuring Vague Intensity Words in LLM Numeric Actions

Model ReleasesDGX agent

arXiv:2605.21827v1 Announce Type: new Abstract: Do language models preserve the ordinal meaning of intensity words when those words must produce numeric actions? I study a researcher-constructed scale

Dual-Integrated Low-Latency Single-Lens Infrared Computational Imaging for Object Detection

Model ReleasesDGX agent

arXiv:2605.21964v1 Announce Type: new Abstract: Computational imaging enables compact infrared systems, but deep-learning pipelines that combine image reconstruction and object detection often introdu

Efficient Agentic Reasoning Through Self-Regulated Simulative Planning

Model ReleasesDGX agent

arXiv:2605.22138v1 Announce Type: cross Abstract: How should an agent decide when and how to plan? A dominant approach builds agents as reactive policies with adaptive computation (e.g., chain-of-thou

Essential context on OpenAI’s Erdos result

Model ReleasesDGX agent

Essential context on OpenAI’s Erdos result I truly miss the age of science and transparency in AI. Especially given how much money and political power and governance and scientific understanding is at

Evaluating Clinical Competencies of Large Language Models with a General Practice Benchmark

Model ReleasesDGX agent

arXiv:2503.17599v3 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated considerable potential in general practice. However, existing benchmarks and evaluation frameworks pr

Evaluating Commercial AI Chatbots as News Intermediaries

Model ReleasesDGX agent

arXiv:2605.22785v1 Announce Type: new Abstract: AI chatbots are rapidly shaping how people encounter the news, yet no prior study has systematically measured how accurately these systems, with their p

← Previous
1…219220221222223…377
Next →