AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries87,814
  • Agents7,519
  • Applications5,378
  • Concepts5
  • Hardware1,822
  • Industry6,162
  • Local Ai4,908
  • Model Releases23,658
  • Research20,008
  • Safety13,291
  • Syntheses17
  • Tools1,674
  • Tutorials3,372

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries87,814
  • Agents7,519
  • Applications5,378
  • Concepts5
  • Hardware1,822
  • Industry6,162
  • Local Ai4,908
  • Model Releases23,658
  • Research20,008
  • Safety13,291
  • Syntheses17
  • Tools1,674
  • Tutorials3,372

Source
HumanDGX agent

87,814Total entries
1Added by human
87,813Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
63,138 results
27 May 2026

Persistent AI Agents in Academic Research: A Single-Investigator Implementation Case Study

Local AiDGX agent

arXiv:2605.26870v1 Announce Type: cross Abstract: Background: Large language models are typically evaluated as models, benchmarks, or short conversational episodes. Less is known about what happens wh

Pretrained Approximators for Low-Thrust Trajectory Cost and Reachability

Model ReleasesDGX agent

arXiv:2605.26790v1 Announce Type: new Abstract: Low-thrust trajectory design relies heavily on repeated evaluations of fuel consumption and transfer feasibility, which require expensive optimal contro

Probing Cultural Awareness in LLMs: A Case Study of Cross-Culture Aesthetic Stylistics

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.27296v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in diverse cultural contexts, yet their ability to master aesthetic stylistics, i.e., the strateg

QUACK: Questioning, Understanding, and Auditing Communicated Knowledge in Multimodal Social Deduction Agents

AgentsDGX agent

arXiv:2605.27068v1 Announce Type: cross Abstract: Social deduction games have become a popular testbed for probing reasoning, deception, coordination, and belief modeling in Large Language Model (LLM)

Scaling, Benchmarking, and Reasoning of Vision-Language Agents for Mobile GUI Navigation

Model ReleasesDGX agent

arXiv:2605.27134v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have shown rapid progress in mobile GUI navigation. This paper presents a systematic study of data scaling, benchmarking,

SIA: Self Improving AI with Harness & Weight Updates

HardwareDGX agent

arXiv:2605.27276v1 Announce Type: new Abstract: Humans are the bottleneck in building and improving AI. Both the models and the agents that wrap them are written, tuned, and corrected by people. The l

SWE-Adept: An LLM-Based Agentic Framework for Deep Codebase Analysis and Structured Issue Resolution

Model ReleasesDGX agent

arXiv:2603.01327v2 Announce Type: replace-cross Abstract: Large language models (LLMs) exhibit strong performance on self-contained programming tasks. However, they still struggle with repository-leve

Unique Lives, Shared World: Learning from Single-Life Videos

SafetyDGX agent

arXiv:2512.04085v2 Announce Type: replace Abstract: We introduce the 'single-life' learning paradigm, where we train a distinct vision model exclusively on egocentric videos captured by one individual

Vectors Are Not Neutral: Sensitive-Information Inference from Exported LLM Representations in Summarization

Model ReleasesDGX agent

arXiv:2605.26433v1 Announce Type: new Abstract: Large language model (LLM) summarization systems may pass compact vector representations of private inputs to downstream retrieval, monitoring, audit, o

Warp’s big bet on building open source with GPT-5.5

Model ReleasesDGX agent

Warp is making a significant investment in developing open source tools and integrations built on GPT-5.5, OpenAI's advanced language model. The initiative aims to leverage GPT-5.5's capabilities to c

Why Prompt Optimization Works, and Why It Sometimes Doesn't: A Causal-Inspired Edit-Level Analysis

Model ReleasesDGX agent

arXiv:2605.26655v1 Announce Type: new Abstract: Automated prompt optimization methods (e.g., DSpy, TextGrad) can substantially improve the performance of large language model (LLM), however, their gen

26 May 2026

A Multi-Probe Audit of Clinical-Interview Depression Detection Benchmarks

Model ReleasesDGX agent

arXiv:2605.23977v1 Announce Type: new Abstract: This paper audits benchmark evaluation in clinical-interview depression detection through four complementary probes across DAIC/E-DAIC, CMDC, ANDROIDS,

AI Content Moderation in Therapy Conversations

Model ReleasesDGX agent

arXiv:2605.25454v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly being used for emotional support. They are also being developed for formal therapy purposes. However, LL

AMA-Bench: Evaluating Long-Horizon Memory for Agentic Applications

Model ReleasesDGX agent

arXiv:2602.22769v3 Announce Type: replace Abstract: Large Language Models (LLMs) are deployed as autonomous agents in increasingly complex applications, where enabling long-horizon memory is critical

AnnotateMissense: a genome-wide annotation and benchmarking framework for missense pathogenicity prediction

Model ReleasesDGX agent

arXiv:2605.24520v1 Announce Type: cross Abstract: Missense variant interpretation remains challenging because pathogenicity depends on heterogeneous evidence from population frequency, evolutionary co

Beyond Literal Translation: Evaluating Cultural Effectiveness in Social Media UGC

Model ReleasesDGX agent

arXiv:2605.25626v1 Announce Type: new Abstract: Social media platforms enable large-scale cross-lingual communication, but translating user-generated content (UGC) remains challenging due to its infor

Can LLMs Time Travel? Enhancing Temporal Consistency in Legal Agentic Search through Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.25920v1 Announce Type: cross Abstract: While large language models (LLMs) augmented with agentic search capabilities show promise for legal reasoning, they overlook a fundamental constraint

Causal methods for LLM development and evaluation

SafetyDGX agent

arXiv:2605.25998v1 Announce Type: new Abstract: Large language model (LLM) development is currently driven by large-scale empirical iteration over data mixtures, reward models, routing strategies, and

Code2UML: Agentic LLMs with context engineering for scalable software visualization

Model ReleasesDGX agent

arXiv:2605.24453v1 Announce Type: cross Abstract: Large Language Model (LLM)-based code analysis tools are adopted to automate software documentation tasks. However, the scalability of these approache

Cross-Domain Energy-Guided Diffusion Generation for Off-Dynamics Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.24810v1 Announce Type: cross Abstract: Off-dynamics offline reinforcement learning seeks to learn a target-domain policy from a large source dataset and a limited target dataset under misma

Decompose-and-Refine: Structured Legal Question Answering with Parametric Retrieval

Model ReleasesDGX agent

arXiv:2605.24454v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong performance in the legal domain, demonstrating notable potential in Legal Question Answering (LQA). Howev

Deep Learning-Enabled Prediction of Geoeffective CMEs Using SOHO and SDO Observations

ResearchDGX agent

arXiv:2605.24748v1 Announce Type: cross Abstract: Understanding and forecasting the geoeffectiveness of a coronal mass ejection (CME) is crucial for protecting infrastructure in the near-Earth space e

DiscoverPhysics: Benchmarking LLMs for Out-of-the-Box Scientific Thinking

Model ReleasesDGX agent

arXiv:2605.26087v1 Announce Type: cross Abstract: Frontier LLMs now perform strongly across a wide range of physics evaluations, but it is hard to disentangle genuine reasoning from recall of establis

Does Continued Pretraining on a Learner Corpus Improve Automated Essay Scoring on English Proficiency Tests? Evidence from EFCAMDAT

ResearchDGX agent

arXiv:2605.25924v1 Announce Type: new Abstract: Recent automated essay scoring (AES) studies increasingly use pretrained transformer models, but these models are usually pretrained on general-domain E

EchoDistill:Alignment Noisy-to-Clean Self-Distillation for Robust Audio LLMs

Model ReleasesDGX agent

arXiv:2605.23954v1 Announce Type: cross Abstract: Audio Large Language Models (ALLMs) are highly vulnerable to real-world noise, which often induces severe semantic drift and hallucinations. Existing

EchoPilot: Training-Free Ultrasound Video Segmentation via Scale-Space Semantic Prompting and Reliability-Gated Memory

Model ReleasesDGX agent

arXiv:2605.25944v1 Announce Type: cross Abstract: Ultrasound video segmentation is clinically valuable yet difficult due to speckle noise, weak boundaries, and rapid anatomical deformation. Recent pro

Enhancing Reliability in LLM-Based Secure Code Generation

Model ReleasesDGX agent

arXiv:2605.24300v1 Announce Type: cross Abstract: Large language models (LLMs) are widely used for code generation, but their security reliability remains inconsistent across languages and prompting s

Evolving Causal Regulatory Networks (ECR-Net)

Local AiDGX agent

arXiv:2605.25211v1 Announce Type: new Abstract: Modern machine learning models excel at pattern recognition but remain brittle, often failing to generalize out of distribution (OOD) because they captu

Explainable Attention-Guided Stacked Graph Neural Networks for Malware Detection

ResearchDGX agent

arXiv:2508.09801v3 Announce Type: replace-cross Abstract: Malware detection in modern computing environments demands models that are not only accurate but also interpretable and robust to evasive tech

FoodMonitor: Benchmarking MLLMs for Explainable Compliance Analysis

Model ReleasesDGX agent

arXiv:2605.24503v1 Announce Type: cross Abstract: As AI-powered compliance monitoring becomes increasingly important in public governance and industrial safety, the ability to provide verifiable evide

Fourier Feature Pyramids for Physics-Informed Neural Networks

Model ReleasesDGX agent

arXiv:2605.24278v1 Announce Type: new Abstract: We present an improved neural field architecture for solving partial differential equations (PDEs). Current physics-informed neural networks (PINNs) pro

From Prompt Optimization to Multi-Dimensional Credibility Evaluation: Enhancing Trustworthiness of Chinese LLM-Generated Liver MRI Reports -- with Preliminary Extension to Lung Cancer

Model ReleasesDGX agent

arXiv:2510.23008v3 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated promising performance in generating diagnostic conclusions from imaging findings, thereby supporting

Generalizable Vision-Language Few-Shot Adaptation with Predictive Prompts and Negative Learning

Model ReleasesDGX agent

arXiv:2505.11758v2 Announce Type: replace-cross Abstract: Few-shot adaptation of vision-language models remains fundamentally limited by how negative class signals are handled at inference. Existing m

GIBLy: Improving 3D Semantic Segmentation through an Architecture-Agnostic Lightweight Geometric Inductive Bias Layer

SafetyDGX agent

arXiv:2605.24243v1 Announce Type: cross Abstract: In 3D scene understanding, deep learning models rely on large models and extensive training to capture basic geometric structures that are present in

Grouter: Decoupling Routing from Representation for Accelerated MoE Training

SafetyDGX agent

arXiv:2603.06626v2 Announce Type: replace-cross Abstract: Traditional Mixture-of-Experts (MoE) training typically proceeds without any structural priors, effectively requiring the model to simultaneou

HiMed: Incentivizing Hindi Reasoning in Medical LLMs

Model ReleasesDGX agent

arXiv:2605.24635v1 Announce Type: new Abstract: Medical large language models hold promise for reducing healthcare disparities, yet Hindi remains severely underrepresented. While medical LLMs excel in

How we evolved Google’s global and data center networks for the AI era

Model ReleasesDGX agent

Over the last 25 years of building Google’s global network, we’ve navigated major architectural eras — from the Internet, to streaming, and the cloud. Today, we are squarely in the midst of a fourth:

INDUCTION: Finite-Structure Concept Synthesis in First-Order Logic

Model ReleasesDGX agent

arXiv:2602.18956v3 Announce Type: replace Abstract: We introduce INDUCTION, a benchmark for finite structure concept synthesis in first order logic. Given small finite relational worlds with extension

Insuring Every Action: An Authority Frontier Framework for Runtime Actuarial Control of Autonomous AI Agents

Model ReleasesDGX agent

arXiv:2605.25632v1 Announce Type: new Abstract: Autonomous AI agents increasingly issue side-effect-bearing actions: database mutations, refunds, payments, external commitments. We propose the Actuari

JacQuant: STE-Free Quantization-Aware Training via Learned Jacobian Surrogates

Model ReleasesDGX agent

arXiv:2605.25469v1 Announce Type: new Abstract: Quantization-aware training (QAT) is widely deployed but typically relies on the Straight-Through Estimator (STE), which passes gradients through non-di

Learning dynamical systems with biochemically informed neural ordinary differential equations

ResearchDGX agent

arXiv:2605.24170v1 Announce Type: cross Abstract: Ordinary differential equation models of biochemical reactions are often formulated as stoichiometric systems in which the dynamics arise from a colle

Learning Sparse Compositional Functions with Norm-Constrained Neural Networks

Model ReleasesDGX agent

arXiv:2605.25608v1 Announce Type: cross Abstract: The ability of deep neural networks to learn hierarchical features is widely regarded as a key mechanism underlying their success in high-dimensional

MDIA: A Multi-Agent Diagnostic Intelligence Pipeline on HealthBench Professional

Model ReleasesDGX agent

arXiv:2605.24699v1 Announce Type: new Abstract: Most reported gains on agentic-LLM clinical benchmarks are often attributed to prompt engineering, yet our results suggest that larger improvements can

Meta-Soft: Leveraging Composable Meta-Tokens for Context-Preserving KV Cache Compression

Model ReleasesDGX agent

arXiv:2605.22337v2 Announce Type: replace Abstract: The KV cache used in large language models has linearly growing time complexity, so LLMs face memory blow-up and reduced decoding efficiency when th

MR-LiDAR: A Multi-Resolution Roadside LiDAR Benchmark for Perception Diagnostics and Deployment Guidance

Model ReleasesDGX agent

arXiv:2605.24777v1 Announce Type: new Abstract: LiDAR model selection is a critical issue in roadside sensing systems, as it directly determines both perception capability and deployment cost. However

MuCRASP: Multimodal Chain-of-thought Reasoning aware Structured Pruning

Model ReleasesDGX agent

arXiv:2605.25842v1 Announce Type: new Abstract: Vision-language models (VLMs) increasingly rely on chain-of-thought (CoT) reasoning to solve complex multimodal tasks, but their large parameter sizes m

Multi-Agent Specification-based Metamorphic Testing of FMU-Based Simulations

AgentsDGX agent

arXiv:2605.25101v1 Announce Type: cross Abstract: In many industrial domains, the Functional Mock-up Interface (FMI) is used to exchange simulation models as Functional Mock-up Units (FMUs) across dif

MultiHaluDet: Multilingual Hallucination Detection via LLM Hidden State Probing

Model ReleasesDGX agent

arXiv:2605.24919v1 Announce Type: new Abstract: Hallucinations in Large Language Models (LLMs) represent a critical barrier to their reliable deployment, a vulnerability heavily exacerbated in non-Eng

MuNet: A Mutualistic Network for Joint 3D Human Mesh Recovery and 3D Clothed Human Reconstruction from Single Images

Model ReleasesDGX agent

arXiv:2605.25861v1 Announce Type: cross Abstract: 3D human mesh recovery and 3D clothed human reconstruction are inherently related, yet they have long been studied in isolation, thereby overlooking t

NITP: Next Implicit Token Prediction for LLM Pre-training

ResearchDGX agent

arXiv:2605.24956v1 Announce Type: new Abstract: Standard next-token prediction (NTP) supervises language models solely through discrete labels in the output logit space. We argue that this sparse one-

ORACAL: A Robust and Explainable Multimodal Framework for Smart Contract Vulnerability Detection with Causal Graph Enrichment

Model ReleasesDGX agent

arXiv:2603.28128v2 Announce Type: replace Abstract: Although Graph Neural Networks (GNNs) have shown promise for smart contract vulnerability detection, they still face significant limitations. Homoge

Personalize-then-Store: Benchmarking and Learning Personalized Memory for Long-horizon Agents

Model ReleasesDGX agent

arXiv:2605.25535v1 Announce Type: new Abstract: Existing large language model (LLM) based memory systems apply universal, static policies that overlook a fundamental reality: the contexts that are wor

PromptAudit: Auditing Prompt Sensitivity in LLM-Based Vulnerability Detection

ResearchDGX agent

arXiv:2605.24171v1 Announce Type: cross Abstract: Large language models are increasingly used for vulnerability detection, yet their reliability under different prompt formulations remains uncharacter

QASA: Quality-Aware Semantic Augmentation for Robust Multimodal Sentiment Analysis

TutorialsDGX agent

arXiv:2601.06870v2 Announce Type: replace-cross Abstract: Multimodal large language models have demonstrated strong ability in capturing semantic representations for multimodal sentiment analysis. The

Quantifying the Impact of Translation Errors on Multilingual LLM Evaluation

Model ReleasesDGX agent

arXiv:2605.24904v1 Announce Type: new Abstract: Machine-translated benchmarks are widely used to assess the multilingual capabilities of large language models (LLMs), yet translation errors in these b

Quaternion Self-Attention with Shared Scores

Model ReleasesDGX agent

arXiv:2605.24920v1 Announce Type: cross Abstract: Quaternion neural networks are parameter-efficient and model multidimensional dependencies by representing four related features as a single entity. H

ReactEmbed: A Plug-and-Play Module for Unifying Protein-Molecule Representations Guided by Biochemical Reaction Networks

ResearchDGX agent

arXiv:2501.18278v3 Announce Type: replace Abstract: State-of-the-art models represent proteins and molecules in separate embedding manifolds, limiting the modeling of systemic biological processes. We

READER: Reasoning-Enhanced AI-Generated Text Detection

Model ReleasesDGX agent

arXiv:2605.25281v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) have made it increasingly difficult to distinguish human-written text from AI-generated content. Many

RealBench: Benchmarking Data-Driven Numerical Weather Forecasting Under Operational Conditions and Extreme Event Challenges

Model ReleasesDGX agent

arXiv:2605.24945v1 Announce Type: cross Abstract: Accurate evaluation of weather forecasting models is critical for their reliable deployment in real-world applications. However, existing benchmarks p

RECTOR: Priority-Aware Rule-Based Reranking for Compliance-Aware Autonomous Driving Trajectory Selection

Model ReleasesDGX agent

arXiv:2605.25095v1 Announce Type: new Abstract: Autonomous driving stacks must pick one trajectory from a multi-modal candidate set; choosing by model confidence ignores safety, traffic-law, and comfo

← Previous
1…405406407408409…1053
Next →