AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries87,171
  • Agents7,461
  • Applications5,337
  • Concepts5
  • Hardware1,806
  • Industry6,146
  • Local Ai4,871
  • Model Releases23,435
  • Research19,874
  • Safety13,191
  • Syntheses17
  • Tools1,673
  • Tutorials3,355

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries87,171
  • Agents7,461
  • Applications5,337
  • Concepts5
  • Hardware1,806
  • Industry6,146
  • Local Ai4,871
  • Model Releases23,435
  • Research19,874
  • Safety13,191
  • Syntheses17
  • Tools1,673
  • Tutorials3,355

Source
HumanDGX agent

87,171Total entries
1Added by human
87,170Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,612 results
29 May 2026

When and How Long? The Readout-Mediator Angle in Temporal Reasoning

SafetyDGX agent

arXiv:2605.29126v1 Announce Type: cross Abstract: A linear probe can decode a representation almost perfectly and yet be completely irrelevant to how the model uses it. On calendar-date duration reaso

28 May 2026

Adaptive Reservoir Computing for Multi-Scenario Chaotic System Forecasting

Model ReleasesDGX agent

arXiv:2605.28145v1 Announce Type: new Abstract: We present an adaptive reservoir computing framework for the CTF-4-Science Lorenz benchmark, which evaluates machine learning models across twelve disti

AndroidDaily: A Verifiable Benchmark for Mobile GUI Agents on Real-World Closed-Source Applications

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases
DGX agent

arXiv:2605.27761v1 Announce Type: new Abstract: The rapid development of GUI foundation models and mobile GUI agents has spurred numerous evaluation benchmarks, yet most rely on simulated environments

Beyond being fast, LiteParse is designed to provide highly accurate, semantically coherent text for LLM use. We benchmarked every open-sourc…

Model ReleasesDGX agent

Beyond being fast, LiteParse is designed to provide highly accurate, semantically coherent text for LLM use. We benchmarked every open-source, model-free PDF parser on LLM QA tasks - from PyPDF to PyM

Camellia: Benchmarking Cultural Biases in LLMs for Asian Languages

Model ReleasesDGX agent

arXiv:2510.05291v2 Announce Type: replace Abstract: As Large Language Models (LLMs) develop stronger multilingual capabilities, their sensitivity to culturally diverse entities becomes increasingly im

ClinicalAgents: Multi-Agent Orchestration for Clinical Decision Making with Dual-Memory

Model ReleasesDGX agent

arXiv:2603.26182v2 Announce Type: replace Abstract: While Large Language Models (LLMs) have demonstrated potential in healthcare, they often struggle with the complex, non-linear reasoning required fo

Cost-Sensitive Evaluation for Binary Classifiers

Model ReleasesDGX agent

arXiv:2510.22016v2 Announce Type: replace Abstract: Selecting an appropriate evaluation metric for classifiers is crucial for model comparison, parameter optimization, and deployment decisions, yet th

DeepSciVerify: Verifying Scientific Claim--Citation Alignment via LLM-Driven Evidence Escalation

Model ReleasesDGX agent

arXiv:2605.27710v1 Announce Type: new Abstract: Misalignment between claims and their cited evidence is a common failure mode in reports generated by large language models, limiting their reliability

Diffusion-Based Ukrainian Handwritten Text Generation with Cross-Domain Style Transfer

Model ReleasesDGX agent

arXiv:2605.27487v1 Announce Type: cross Abstract: Handwritten text generation (HTG) conditioned on writer style has been widely studied for Latin scripts, but remains underexplored for low-resource an

Do Audio LLMs Listen or Read? Analyzing and Mitigating Paralinguistic Failures with VoxParadox

Model ReleasesDGX agent

arXiv:2605.27772v1 Announce Type: cross Abstract: Audio large language models (Audio LLMs) demonstrate strong performance on speech understanding tasks, yet their ability to understand paralinguistic

EAGer: Entropy-Aware GEneRation for Adaptive Inference-Time Scaling

ResearchDGX agent

arXiv:2510.11170v2 Announce Type: replace-cross Abstract: With the rise of reasoning language models and test-time scaling methods as a paradigm for improving model performance, substantial computatio

Efficient Pre-Training of LLMs through Truncated SVD Layers

Model ReleasesDGX agent

arXiv:2605.28573v1 Announce Type: cross Abstract: The massive scaling of Large Language Models (LLMs) has made pretraining increasingly cost-prohibitive. While low-rank representation and orthonormal

Escape the Language Prior: Mitigating Late-Stage Modality Collapse in Audio Reasoning via Modality-Aware Policy Optimization

SafetyDGX agent

arXiv:2605.27741v1 Announce Type: new Abstract: Audio and omni-modal large language models exhibit impressive cross-modal reasoning capabilities. However, applying standard reinforcement learning post

Evolving Dataflow to process massive datasets for machine learning

Model ReleasesDGX agent

Google created MapReduce more than 20 years ago to solve the scaling problems in data processing that the then young company was running into. The AI era that we are in now demands efficient, large-sc

Explaining is Harder Than Predicting Alone: Evaluating Concept-based Explanations of MLLMs as ICL Visual Classifiers

ResearchDGX agent

arXiv:2605.28215v1 Announce Type: new Abstract: In-context learning (ICL) enables multimodal large language models (MLLMs) to classify images from a few labelled examples. Yet, how these models use th

HELEA: Hard-Negative Benchmark and LLM-based Reranking for Robust Entity Alignment

Model ReleasesDGX agent

arXiv:2605.28308v1 Announce Type: new Abstract: Entity Alignment (EA) is essential for knowledge graph (KG) fusion, but existing benchmarks often allow models to exploit name overlap rather than relat

IAR2: Improving Autoregressive Visual Generation with Semantic-Detail Associated Token Prediction

Model ReleasesDGX agent

arXiv:2510.06928v2 Announce Type: replace Abstract: Autoregressive models have emerged as a powerful paradigm for visual content creation, but often overlook the intrinsic structural properties of vis

IFMTBench: A Comprehensive Benchmark for Multilingual Translation Instruction Following

Model ReleasesDGX agent

arXiv:2605.28218v1 Announce Type: new Abstract: Modern translation workflows demand more than semantic equivalence. Users routinely require models to preserve JSON or HTML schemas, honor curated gloss

Inpainting-Style Conditional Diffusion for Multivariable Time Series Forecasting

Model ReleasesDGX agent

arXiv:2605.28324v1 Announce Type: new Abstract: In this paper, we propose a novel conditional diffusion-based framework for multivariable time-series solar power forecasting. The proposed method refor

Knowing When to Ask: Segment-Level Credit Assignment for LLM Tool Use

AgentsDGX agent

arXiv:2605.27788v1 Announce Type: cross Abstract: Humans know when to reach for help e.g. 347 imes 28 warrants a calculator while 2+2 does not. Language models do not. Prompt-based approaches can inst

Mags-RL: Wearing Multimodal LLMs a Magnifying Glass via Agentic Reinforcement Learning For Complex Scene Reasoning

Model ReleasesDGX agent

arXiv:2605.27960v1 Announce Type: new Abstract: Despite their popularity and success, Multimodal Large Language Models (MLLMs) often struggle to interpret images accurately, which limits their reasoni

NL-MambaXCT: Self-Supervised Nested-Learning Mamba for Nomex Honeycomb X-ray CT Defect Classification

Model ReleasesDGX agent

arXiv:2605.27454v1 Announce Type: cross Abstract: X-ray computed tomography (XCT) is widely used for non-destructive testing of Nomex honeycomb structures in aerospace manufacturing, but industrial in

Patched-DeltaNet: Token-Level Event-Driven Memory for Linear-Time Anomaly Detection

Model ReleasesDGX agent

arXiv:2605.27992v1 Announce Type: new Abstract: Time series anomaly detection is critical for maintaining the reliability of mission-critical systems. While Transformer-based models like PatchTST have

Random Process Flow Matching: Generative Implicit Representations of Multivariate Random Fields

TutorialsDGX agent

arXiv:2605.28625v1 Announce Type: new Abstract: Generative modeling provides a powerful framework for learning data distributions. These models initially relied on probabilistic methods such as Gaussi

RULER: Representation-Level Verification of Machine Unlearning

ResearchDGX agent

arXiv:2605.27569v1 Announce Type: new Abstract: Machine unlearning aims to remove the influence of specific training records from a deployed model without retraining from scratch. Current protocols ve

Satisfiability Solving with LLMs: A Matched-Pair Evaluation of Reasoning Capability

ResearchDGX agent

arXiv:2605.28602v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for tasks that implicitly reduce to Boolean satisfiability (SAT), yet their reasoning ability on SAT

SelfJudge: Faster Speculative Decoding via Self-Supervised Judge Verification

ResearchDGX agent

arXiv:2510.02329v2 Announce Type: replace-cross Abstract: Speculative decoding accelerates LLM inference by verifying candidate tokens from a draft model against a larger target model. Recent judge de

Simorgh at SemEval-2026 task 7: Region-Aware Hybrid Retrieval for Low-Resource Cultural Reasoning in Multilingual Question Answering

Model ReleasesDGX agent

arXiv:2605.27636v1 Announce Type: new Abstract: Although Large Language Models (LLMs) demonstrate excellent capabilities and performance for general reasoning tasks within the general public domain, t

SmartIterator: Visual Analytics Workflows for Supervising Unsupervised Data Grouping

Model ReleasesDGX agent

arXiv:2605.28219v1 Announce Type: cross Abstract: Unsupervised learning methods -- topic modeling, partition-based and density-based clustering -- produce data groupings without human guidance, yet ch

Soft Specialists: alpha-Renyi Ensembles for Uncertainty-Aware LLM Post-Training

Local AiDGX agent

arXiv:2605.27747v1 Announce Type: cross Abstract: Existing training approaches for large language models learn a single set of parameters, based on large volumes of data, which is typically heterogene

Towards Reliable Multilingual LLMs-as-a-Judge: An Empirical Study

ResearchDGX agent

arXiv:2605.28710v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for the automatic evaluation of generated text, yet most prior work focuses on English. Despite the

Universal Time Series Generation with Neural Controlled Differential Equations

ResearchDGX agent

arXiv:2605.28507v1 Announce Type: new Abstract: Recent work on the sequence universality of State Space Models (SSMs) has introduced efficient, maximally expressive continuous-time approaches for time

VULPO: Context-Aware Vulnerability Detection via On-Policy LLM Optimization

Model ReleasesDGX agent

arXiv:2511.11896v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have recently shown strong potential in vulnerability detection (VD). However, accurately detecting vulnerabiliti

WeatherCity: Urban Scene Reconstruction with Controllable Multi-Weather Transformation

Model ReleasesDGX agent

arXiv:2602.22096v2 Announce Type: replace Abstract: Editable high-fidelity 4D scenes are crucial for autonomous driving, as they can be applied to end-to-end training and closed-loop simulation. Howev

27 May 2026

Axial-Centric Cross-Plane Attention for 3D Medical Image Classification

Model ReleasesDGX agent

arXiv:2602.21636v2 Announce Type: replace Abstract: Abridged: Clinicians commonly interpret 3D medical images by examining multiple anatomical planes rather than relying on volumetric views. In clinic

Causal Representation Learning for Generalisable Recommendation

Model ReleasesDGX agent

arXiv:2605.27043v1 Announce Type: cross Abstract: Predictive models trained on observational data often fail to generalise to the distributions they encounter when deployed, especially when the traini

Drive-P2D: A Progressive Perception-to-Decision Benchmark for VLMs in Autonomous Driving

Model ReleasesDGX agent

arXiv:2601.14702v2 Announce Type: replace Abstract: Autonomous driving requires reliable perception and safe decision-making in complex scenarios. Recent vision-language models (VLMs) demonstrate reas

ECHO-2: A Large-Scale Distributed Rollout Framework for Cost-Efficient Reinforcement Learning

Model ReleasesDGX agent

arXiv:2602.02192v5 Announce Type: replace Abstract: Reinforcement learning (RL) is a critical stage in post-training large language models (LLMs), involving repeated interaction between rollout genera

EpiCurveBench: Evaluating VLMs on Epidemic Curve Digitization

Model ReleasesDGX agent

arXiv:2605.27195v1 Announce Type: new Abstract: Chart-to-data extraction with vision-language models (VLMs) is increasingly evaluated on benchmarks that show diminishing headroom (frontier VLMs exceed

Error Analysis of Discrete Flow with Generator Matching

ResearchDGX agent

arXiv:2509.21906v3 Announce Type: replace-cross Abstract: Discrete flow models offer a powerful framework for learning distributions over discrete state spaces and have demonstrated superior performan

InfoSynth: Information-Guided Benchmark Synthesis for LLMs

Model ReleasesDGX agent

arXiv:2601.00575v2 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated significant advancements in reasoning and code generation, but efficiently creating new benchmarks to

LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness

Model ReleasesDGX agent

arXiv:2605.26438v1 Announce Type: cross Abstract: Large language models can recognize when they are being evaluated (evaluation awareness) and behave differently because of that, which undermines the

Med-CoReasoner: Reducing Language Disparities in Medical Reasoning via Language-Informed Co-Reasoning

Model ReleasesDGX agent

arXiv:2601.08267v3 Announce Type: replace Abstract: While reasoning-enhanced large language models perform strongly on English medical tasks, a persistent multilingual gap remains, with substantially

MobileMoE: Scaling On-Device Mixture of Experts

Model ReleasesDGX agent

arXiv:2605.27358v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) has become the de facto architecture for hundred-billion-parameter language models, yet its advantages at sub-billion scales

More Expressive Feedforward Layers: Part I. Token-Adaptive Mixing of Activations

Model ReleasesDGX agent

arXiv:2605.26647v1 Announce Type: cross Abstract: Feedforward network (FFN) layers account for a large fraction of parameters and nonlinear expressivity in Transformer-based large language models (LLM

MSCGC-KAN: Multi-scale Causal Graph Convolution and Kolmogorov-Arnold Feature Mapping for EEG Emotion Recognition

ResearchDGX agent

arXiv:2605.26624v1 Announce Type: new Abstract: Electroencephalogram (EEG)-based emotion recognition is an important affective computing task, and recent EEG foundation models provide useful generic r

NestedKV: Nested Memory Routing for Long-Context KV Cache Compression

Model ReleasesDGX agent

arXiv:2605.26678v1 Announce Type: new Abstract: Long-context language models are limited by the memory footprint of the key-value (KV) cache. Existing training-free KV compression methods usually rank

O-MARC: Omni Memory-Augmented Compression Distillation for Efficient Video Understanding

Model ReleasesDGX agent

arXiv:2605.26584v1 Announce Type: new Abstract: Omnimodal large language models enable unified audio video understanding, but long joint token sequences make inference costly, and existing benchmarks

OmniGF: A Dual-Branch Vision-Language Framework for Unified Gaze Following

Local AiDGX agent

arXiv:2605.26399v1 Announce Type: new Abstract: Understanding human gaze behavior is essential for complex scene comprehension and human-computer interaction. Traditional gaze following models are typ

Optimising Factual Consistency in Summarisation via Preference Learning from Multiple Imperfect Metrics

TutorialsDGX agent

arXiv:2605.26840v1 Announce Type: new Abstract: Reinforcement learning with evaluation metrics as rewards is widely used to enhance specific capabilities of language models. However, for tasks such as

PaTAS: A Framework for Trust Propagation in Neural Networks Using Subjective Logic

Model ReleasesDGX agent

arXiv:2511.20586v4 Announce Type: replace Abstract: Trustworthiness has become a key requirement for the deployment of artificial intelligence systems in safety-critical applications. Conventional eva

Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User History

Model ReleasesDGX agent

arXiv:2602.17003v2 Announce Type: replace-cross Abstract: Large language models have advanced web agents, yet current agents lack personalization capabilities. Since users rarely specify every detail

Position: AI Safety Requires Effective Controllability

Model ReleasesDGX agent

arXiv:2605.27117v1 Announce Type: new Abstract: AI safety is still largely framed as alignment: training models to follow human preferences, safety policies, and normative constraints. That framing ha

Probing the Knowledge Boundary: An Interactive Agentic Framework for Deep Knowledge Extraction

AgentsDGX agent

arXiv:2602.00959v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) can be seen as compressed knowledge bases, but it remains unclear what knowledge they truly contain and how far t

SEAL: Self-Evolving Agentic Learning for Conversational Question Answering over Knowledge Graphs

Model ReleasesDGX agent

arXiv:2512.04868v2 Announce Type: replace-cross Abstract: Knowledge-based conversational question answering (KBCQA) confronts persistent challenges in resolving coreference, modeling contextual depend

SONAR-LLM: Autoregressive Transformer that Thinks in Sentence Embeddings and Speaks in Tokens

Model ReleasesDGX agent

arXiv:2508.05305v2 Announce Type: replace Abstract: The recently proposed Large Concept Model (LCM) generates text by predicting a sequence of sentence-level embeddings and training with either mean-s

Strategies for Guiding LLMs to Use Software Design Patterns: A Case of Singleton

Model ReleasesDGX agent

arXiv:2605.26898v1 Announce Type: cross Abstract: Large Language Models (LLMs) can generate functional source code from natural-language prompts, but often fail to consistently follow higher-level arc

Structured Relational Reasoning for Group Activity Assessment

Model ReleasesDGX agent

arXiv:2508.07996v2 Announce Type: replace Abstract: Group Activity Detection (GAD) involves recognizing social groups and their collective behaviors in videos. Vision Foundation Models (VFMs), like DI

TowerMind: A Tower Defence Game Learning Environment and Benchmark for LLM as Agents

Model ReleasesDGX agent

arXiv:2601.05899v2 Announce Type: replace Abstract: Recent breakthroughs in Large Language Models (LLMs) have positioned them as a promising paradigm for agents, with long-term planning and decision-m

What Makes Chain-of-Thought Work at Probe Time? Local Co-occurrence Rather Than Global Derivation

Model ReleasesDGX agent

arXiv:2605.26795v1 Announce Type: new Abstract: Chain-of-thought (CoT) prompting reliably improves language-model accuracy, but which properties of a rationale text drive the improvement is poorly und

← Previous
1…348349350351352…1044
Next →