AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,619Total entries
1Added by human
84,618Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,603 results
26 May 2026

'Si'multaneous 'S'patial-'T'emporal Message Passing for Dynamic Graph Representation Learning

Model ReleasesDGX agent

arXiv:2605.25548v1 Announce Type: cross Abstract: Dynamic graph neural networks (DGNNs) that operate on snapshot sequences typically fall into one of two categories. Temporal-first approaches build pe

SimuWoB: Simulating Real-World Mobile Apps for Fast and Faithful GUI Agent Benchmarking

Model ReleasesDGX agent

arXiv:2605.25160v1 Announce Type: new Abstract: Mobile GUI agents powered by large language models have progressed rapidly, creating urgent needs for realistic and comprehensive evaluation. Existing b

SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills

Model Releases

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2605.24117v1 Announce Type: new Abstract: Large language model (LLM) agents accumulate rich episodic trajectories while solving real-world tasks, but it remains unclear whether such experience c

SLAP: Stratified Loss-based Pruning for On-Policy Data-Efficient Instruction Tuning

Model ReleasesDGX agent

arXiv:2605.23969v1 Announce Type: new Abstract: Instruction tuning has optimized the specialized capabilities of large language models (LLMs), but it often requires extensive datasets and prolonged tr

Small Models, Strong Priors: Architectural Inductive Bias for Parameter-Efficient Neural PDE Solvers

Model ReleasesDGX agent

arXiv:2605.25949v1 Announce Type: cross Abstract: Neural PDE solvers have followed the scaling trajectory of vision and language, with recent foundation models reaching billions of parameters. We argu

Smart Timing for Mining: A Deep Learning Framework for Bitcoin Hardware ROI Prediction

Model ReleasesDGX agent

arXiv:2512.05402v2 Announce Type: replace-cross Abstract: Bitcoin mining hardware acquisition requires strategic timing due to volatile markets, rapid technological obsolescence, and protocol-driven r

SMDD-Bench: Can LLMs Solve Real-World Small Molecule Drug Design Tasks?

Model ReleasesDGX agent

arXiv:2605.21740v2 Announce Type: replace Abstract: LLM agents have incredible potential for scientific discovery applications. However, the performance of LLM agents on real-world, small molecule dru

SODE: Analyzing Social Dynamics in LLM Agents

Model ReleasesDGX agent

arXiv:2605.23949v1 Announce Type: cross Abstract: As Large Language Models (LLMs) evolve into interactive agents, understanding their behavioral alignment within human social dynamics becomes essentia

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models

Model ReleasesDGX agent

arXiv:2506.18543v2 Announce Type: replace-cross Abstract: The rapid proliferation of Large Language Models (LLMs) has heightened concerns regarding their exposure to jailbreak attacks, which craft adv

SomaliBench Eval: Measuring English-to-Somali Refusal Gaps in Open-Weight Language Models

Model ReleasesDGX agent

arXiv:2605.25420v1 Announce Type: cross Abstract: Large language model safety evaluation remains heavily English-centered, leaving low-resource languages under-measured even when models are deployed g

Some ideas for what comes next, May 2026

Model ReleasesDGX agent

This article from Interconnects explores speculative ideas and predictions about technological developments and trends expected around May 2026, likely covering emerging AI capabilities, infrastructur

Sources: Chinese government agencies begin imposing overseas travel restrictions on individuals involved in advanced AI work, including at Alibaba and DeepSeek (Bloomberg)

Model ReleasesDGX agent

Bloomberg: Sources: Chinese government agencies begin imposing overseas travel restrictions on individuals involved in advanced AI work, including at Alibaba and DeepSeek — China is restricting overse

Spectral Probe-Circuits: A Three-Step Recipe for Identifying Attention-Head Circuits in Pretrained Transformers

Model ReleasesDGX agent

arXiv:2605.24059v1 Announce Type: cross Abstract: We present a three-step recipe for identifying attention-head circuits in pretrained transformers. A per-head spectral signal -- the time-integrated p

Spectral Retrieval: Multi-Scale Sinc Convolution over Token Embeddings for Localized Retrieval in LLM Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2605.24764v1 Announce Type: cross Abstract: [Abridged] - Spectral Retrieval is a plug-in re-ranking stage that interpolates between per-token MaxSim and mean-pool retrieval through a multi-scale

Split-Merge: A Difference-based Approach for Dominant Eigenvalue Problem

Model ReleasesDGX agent

arXiv:2501.15131v3 Announce Type: replace-cross Abstract: The computation of the dominant eigenpair for symmetric positive semidefinite matrices is fundamental in numerical optimization. This work shi

Spotify launches a library of over 650 narrated long-form magazine articles in English for Premium users; free users can buy articles 'individually for $1.99' (Jess Weatherbed/The Verge)

Model ReleasesDGX agent

Jess Weatherbed / The Verge: Spotify launches a library of over 650 narrated long-form magazine articles in English for Premium users; free users can buy articles “individually for $1.99” — More than

Stein Variational Ergodic Surface Coverage with SE(3) Constraints

Model ReleasesDGX agent

arXiv:2603.09458v3 Announce Type: replace Abstract: Surface manipulation tasks require robots to generate trajectories that comprehensively cover complex 3D surfaces while maintaining precise end-effe

Stochastic Estimation of the Layer-wise Hessian Trace for Monitoring Neural-network Training

Model ReleasesDGX agent

arXiv:2605.25674v1 Announce Type: new Abstract: The loss and the norm of its gradient separate the healthy and the pathological regimes of neural-network training only weakly, whilst the curvature of

Stochastic Linear Bandits with Parameter Noise

Model ReleasesDGX agent

arXiv:2601.23164v2 Announce Type: replace Abstract: We study the stochastic linear bandits with parameter noise model, in which the reward of action a is a^op heta where heta is sampled i.i.d. We show

STREAM: A Data-Centric Framework for Mining High-Value Task-Oriented Dialogues from Streaming Media

Model ReleasesDGX agent

arXiv:2605.25162v1 Announce Type: cross Abstract: Large language models for vertical domains are bottlenecked by the scarcity of complex, domain-specific task-oriented dialogues. Existing data acquisi

Streaming Reinforcement Learning under Partial Observability with Real-Time Recurrent Learning

Model ReleasesDGX agent

arXiv:2605.24709v1 Announce Type: new Abstract: Streaming reinforcement learning has emerged as an online learning paradigm that conforms to the restrictions of natural learning agents that process da

StreamProfileBench: A Benchmark for Fine-Grained User Profile Inference in Real-World Streaming Scenarios

Model ReleasesDGX agent

arXiv:2605.25758v1 Announce Type: new Abstract: Large Language Models (LLMs) have reshaped user profiling, yet current evaluations mainly focus on static data snapshots. This paradigm overlooks the re

StructBreak: Structural Cognitive Overload-Induced Safety Failures in MLLMs

Model ReleasesDGX agent

arXiv:2605.25534v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) excel at structural reasoning yet suffer from a sharp logical brittleness in structural consistency. We term th

Structural Abstraction as an Inductive Bias for Non-Stationary Language Model Training

Model ReleasesDGX agent

arXiv:2603.17198v2 Announce Type: replace-cross Abstract: A foundational principle in cognitive science holds that intelligent agents do not learn by storing experiences as isolated instances, but by

SURGE: On the Potential of Large Language Models as General-Purpose Surrogate Code Executors

Model ReleasesDGX agent

arXiv:2502.11167v5 Announce Type: replace-cross Abstract: Neural surrogate models are powerful and efficient tools in data mining. Meanwhile, large language models (LLMs) have demonstrated remarkable

System scaling is the next real bottleneck in agentic AI. If you build agent orchestration layers, this is a clean map of where the engineer…

Model ReleasesDGX agent

System scaling is the next real bottleneck in agentic AI. If you build agent orchestration layers, this is a clean map of where the engineering leverage actually sits. The labs own the model. You own

Teaching large language models to reason like expert diagnosticians

Model ReleasesDGX agent

arXiv:2509.12194v2 Announce Type: replace Abstract: Differential diagnosis is an iterative process that integrates patient information with broader medical knowledge. Clinical case series such as the

Teaching Through Analogies: A Modular Pipeline for Educational Analogy Generation

Model ReleasesDGX agent

arXiv:2605.24211v1 Announce Type: cross Abstract: Analogies help learners understand unfamiliar concepts by relating them to known concepts. Despite recent advances, large language models (LLMs) conti

Test-Time Graph Search for Goal-Conditioned Reinforcement Learning

Model ReleasesDGX agent

arXiv:2510.07257v2 Announce Type: replace Abstract: Offline goal-conditioned reinforcement learning (GCRL) often struggles with long-horizon tasks, where errors in value estimation accumulate and prod

Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head Generation

Model ReleasesDGX agent

arXiv:2605.25488v1 Announce Type: cross Abstract: Audio-driven talking-head generation has achieved remarkable progress with recent models such as AniTalker, FLOAT, and Sonic. Despite their success, m

The Age of Curiosity Meets the Age of AI: Benchmarking Child Safety in Large Language Models

Model ReleasesDGX agent

arXiv:2605.25510v1 Announce Type: new Abstract: Children increasingly have access to Large Language Models (LLMs), which may expose them to responses that are developmentally inappropriate or require

the basic trick to using Claude Code for non-technical work is to put a bunch of files in a folder and tell it can write scripts + make HTML

Model ReleasesDGX agent

Claude Code can be used for non-technical work by organizing files in a folder and instructing it to write scripts and create HTML documents. This approach allows users without programming expertise t

The LSCD Benchmark: a Testbed for Diachronic Word Meaning Tasks

Model ReleasesDGX agent

arXiv:2404.00176v3 Announce Type: replace Abstract: Lexical Semantic Change Detection (LSCD) is a complex, lemma-level task, which is usually operationalized based on two subsequently applied usage-le

The Model Is Not the Product: A Dual-Pillar Architecture for Local-First Psychological Coaching

Model ReleasesDGX agent

arXiv:2605.24411v1 Announce Type: new Abstract: Existing language model applications struggle to meet the demand for emotionally oriented support, primarily due to their inability to maintain deep, pe

The Perception-Physics Paradox: Probing Scientific Alignment with TC-Bench

Model ReleasesDGX agent

arXiv:2605.24782v1 Announce Type: new Abstract: While Vision Foundation Models (VFMs) excel at predictive tasks on satellite imagery, their performance can arise from visual correlations rather than u

The Time is Here for Just-in-Time Systems: Challenges and Opportunities

Model ReleasesDGX agent

arXiv:2605.24096v1 Announce Type: cross Abstract: Core systems like key-value stores have historically taken years to build, and are designed to be general so as to amortize cost across deployments, p

this basically destroys the extrapolations that had Anthropic making two trillion dollars a year.

Model ReleasesDGX agent

this basically destroys the extrapolations that had Anthropic making two trillion dollars a year. It's clear that growth for coding tools such as Claude Code has decelerated from the pace it was since

this logic from 2024 held up pretty well considering all that’s changed.

Model ReleasesDGX agent

this logic from 2024 held up pretty well considering all that’s changed. 9 reasons that OpenAI could someday be seen as the WeWork of AI: 👉 Lots of competitors are catching up. 👉 OpenAI has been force

TIAR: Trajectory-Informed Advantage Reweighting for LLM Abstention Learning

Model ReleasesDGX agent

arXiv:2605.25850v1 Announce Type: cross Abstract: This paper investigates large language model (LLM) abstention learning, specifically using ternary reward, which incentivize truthfulness in large lan

TimeSpot: Benchmarking Geo-Temporal Understanding in Vision-Language Models in Real-World Settings

Model ReleasesDGX agent

arXiv:2603.06687v2 Announce Type: replace-cross Abstract: Geo-temporal understanding, the ability to infer location, time, and contextual properties from visual input alone, underpins applications suc

Tiny Brains, Giant Impact: Uncovering the Keystone Neurons of LLM with Just a Few Prompts

Model ReleasesDGX agent

arXiv:2605.24846v1 Announce Type: cross Abstract: Large language models (LLMs) display strong comprehensive abilities, yet the internal mechanisms that support these behaviors remain insufficiently un

ToolRegistry: A Protocol-Agnostic Tool Management Library for Function-Calling LLMs

Model ReleasesDGX agent

arXiv:2507.10593v3 Announce Type: replace-cross Abstract: Every LLM tool call is structurally an RPC -- a function name, JSON arguments, and a serialized result -- yet each protocol (native Python, MC

Topology-Driven Transferability Estimation of Medical Foundation Models for Segmentation

Model ReleasesDGX agent

arXiv:2602.23916v2 Announce Type: replace-cross Abstract: The advent of large-scale self-supervised learning (SSL) has produced a vast zoo of medical foundation models. However, selecting optimal medi

Toward a Benchmark for Controllable Simulation of Imperfect Students with Large Language Models

Model ReleasesDGX agent

arXiv:2605.25601v1 Announce Type: cross Abstract: Teacher education requires deliberate practice with learners who exhibit identifiable strengths, weaknesses, and partial mastery. Large language model

Towards Large Model Feature Coding

Model ReleasesDGX agent

arXiv:2605.24025v1 Announce Type: cross Abstract: Large models have delivered remarkable performance across a wide range of perception and generation tasks, yet practical deployment is increasingly co

TRACE: A taxonomy-grounded synthetic dataset for teaching-program generation and session interpretation in Applied Behavior Analysis

Model ReleasesDGX agent

arXiv:2605.25038v1 Announce Type: new Abstract: Applied Behavior Analysis (ABA) is a clinical discipline whose documentation, teaching programs and multi-session behavioral logs, is formulaic and high

TRACER: A Semantic-Aware Framework for Fine-Grained Contamination Detection in Code LLMs

Model ReleasesDGX agent

arXiv:2605.24079v1 Announce Type: cross Abstract: Data contamination is a known threat to the reliability of model evaluation. However, it remains underexplored in code large language models (LLMs), w

Trade-off Functions for DP-SGD with Subsampling based on Random Shuffling: Tight Upper and Lower Bounds

Model ReleasesDGX agent

arXiv:2605.06259v2 Announce Type: replace Abstract: We derive a tight analysis of the trade-off function for Differentially Private Stochastic Gradient Descent (DP-SGD) with subsampling based on rando

TriVAL: A Tri-Validation Framework for Faithful Automatic Optimization Modeling

Model ReleasesDGX agent

arXiv:2605.23966v1 Announce Type: cross Abstract: Optimization modeling serves as the pivotal bridge between natural-language problem descriptions and optimization solvers, and remains a cornerstone f

Trust but Verify: Prover-Verifier Deliberation for Selective LLM Prediction

Model ReleasesDGX agent

arXiv:2605.25133v1 Announce Type: new Abstract: Reliably knowing when a language model is correct is almost as important as being correct. We introduce prover-verifier deliberation (PVD), an inference

Truthful Online Preference Aggregation for LLM Fine-Tuning in Mobile Crowdsourcing

Model ReleasesDGX agent

arXiv:2605.24052v1 Announce Type: cross Abstract: To better serve users' demands in mobile applications (e.g., navigation), mobile crowdsourcing platforms can iteratively align large language model (L

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering

Model ReleasesDGX agent

arXiv:2605.24703v1 Announce Type: cross Abstract: Large language models (LLMs) and time-series language models (TSLMs) are increasingly applied to time-series question answering (TSQA). Unlike text-on

TTPrint: Evidence-Grounded TTP Extraction via Diverge-then-Converge Verification

Model ReleasesDGX agent

arXiv:2605.25836v1 Announce Type: cross Abstract: Extracting MITRE ATT&CK techniques from cyber threat intelligence (CTI) reports is an open-set, multi-label problem requiring both high recall (not mi

TypedCSIP: Typed Counterfactual Pretraining for Chinese Legislative Conflict Classification

Model ReleasesDGX agent

arXiv:2605.25474v1 Announce Type: new Abstract: TypedCSIP is a typed counterfactual pretraining method for the conflict-classification task of the LCR-CN benchmark (Zhao et al., 2026): given a (superi

Uber president says AI spending is getting ‘harder to justify’

Model ReleasesDGX agent

After reportedly exhausting its annual AI budget just four months into 2026, Uber is now questioning whether it's actually seeing meaningful returns on its investments. In an interview with Rapid Resp

Uncertainty Decomposition via Cyclical SG-MCMC and Soft-label Learning for Subjective NLP

Model ReleasesDGX agent

arXiv:2605.24773v1 Announce Type: new Abstract: Annotator disagreement in emotion classification reflects ambiguity intrinsic to emotion concepts and is essential for predictor-quality assessment in s

Understanding and Mitigating Premature Confidence for Better LLM Reasoning

Model ReleasesDGX agent

arXiv:2605.24396v1 Announce Type: new Abstract: Long chains of thought (CoT) from current language models frequently contain logical gaps and unjustified leaps, limiting the gains from additional test

Understanding Conversational Patterns in Multi-agent Programming: A Case Study on Fibonacci Game Development

Model ReleasesDGX agent

arXiv:2605.24138v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly applied to software engineering (SE), yet their potential for autonomous, role-oriented collaboration re

Uni-DPO: A Unified Paradigm for Dynamic Preference Optimization of LLMs

Model ReleasesDGX agent

arXiv:2506.10054v4 Announce Type: replace-cross Abstract: Direct Preference Optimization (DPO) has emerged as a cornerstone of reinforcement learning from human feedback (RLHF) due to its simplicity a

URS: A Unified Neural Routing Solver for Cross-Problem Zero-Shot Generalization

Model ReleasesDGX agent

arXiv:2509.23413v2 Announce Type: replace Abstract: Multi-task neural routing solvers have emerged as a promising paradigm for their ability to solve multiple vehicle routing problems (VRPs) using a s

← Previous
1…214215216217218…377
Next →