AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries90,259
  • Agents7,702
  • Applications5,506
  • Concepts5
  • Hardware1,896
  • Industry6,187
  • Local Ai5,048
  • Model Releases24,516
  • Research20,616
  • Safety13,635
  • Syntheses17
  • Tools1,677
  • Tutorials3,454

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries90,259
  • Agents7,702
  • Applications5,506
  • Concepts5
  • Hardware1,896
  • Industry6,187
  • Local Ai5,048
  • Model Releases24,516
  • Research20,616
  • Safety13,635
  • Syntheses17
  • Tools1,677
  • Tutorials3,454

Source
HumanDGX agent

Content type
AllBlog
90,259Total entries
1Added by human
90,258Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
65,128 results
Safety

Theoretical Foundations and Effective Algorithms for Policy-Aware Simulator Learning

DGX agent

arXiv:2605.29032v1 Announce Type: new Abstract: Model-based reinforcement learning (MBRL) agents typically learn world models by minimizing predictive loss. However, powerful RL optimizers inevitably

safetyarxiv-cs-lg
29 May 2026
Model Releases
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Tiny but Trusted: Efficient Vision-Language Reasoning for Time-Series Anomaly Detection

DGX agent

arXiv:2605.30344v1 Announce Type: new Abstract: Recent advances in Vision-Language Models (VLMs) have achieved impressive performance across many tasks, yet prior studies report unsatisfactory perform

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Took me a while to figure out what all the ESMFold2 rage was about. At first, the benchmarking data didn't look super remarkable to me but i…

DGX agent

Took me a while to figure out what all the ESMFold2 rage was about. At first, the benchmarking data didn't look super remarkable to me but it turns there are many impressive aspects: - Fully open sour

model-releasesyann-lecun--x
29 May 2026
Model Releases

TRACE: Toulmin-based Reasoning Assessment through Constructive Elements for LLM CoT Evaluation

DGX agent

arXiv:2605.29656v1 Announce Type: new Abstract: Evaluating open-ended outputs from large language models (LLMs) remains challenging due to the absence of ground truth. Existing metrics rely on final-a

model-releasesarxiv-cs-ai
29 May 2026
Safety

When and How Long? The Readout-Mediator Angle in Temporal Reasoning

DGX agent

arXiv:2605.29126v1 Announce Type: cross Abstract: A linear probe can decode a representation almost perfectly and yet be completely irrelevant to how the model uses it. On calendar-date duration reaso

safetyarxiv-cs-ai
29 May 2026
Model Releases

Adaptive Reservoir Computing for Multi-Scenario Chaotic System Forecasting

DGX agent

arXiv:2605.28145v1 Announce Type: new Abstract: We present an adaptive reservoir computing framework for the CTF-4-Science Lorenz benchmark, which evaluates machine learning models across twelve disti

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

AndroidDaily: A Verifiable Benchmark for Mobile GUI Agents on Real-World Closed-Source Applications

DGX agent

arXiv:2605.27761v1 Announce Type: new Abstract: The rapid development of GUI foundation models and mobile GUI agents has spurred numerous evaluation benchmarks, yet most rely on simulated environments

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

Beyond being fast, LiteParse is designed to provide highly accurate, semantically coherent text for LLM use. We benchmarked every open-sourc…

DGX agent

Beyond being fast, LiteParse is designed to provide highly accurate, semantically coherent text for LLM use. We benchmarked every open-source, model-free PDF parser on LLM QA tasks - from PyPDF to PyM

model-releasesjerry-liu--x
28 May 2026
Model Releases

Camellia: Benchmarking Cultural Biases in LLMs for Asian Languages

DGX agent

arXiv:2510.05291v2 Announce Type: replace Abstract: As Large Language Models (LLMs) develop stronger multilingual capabilities, their sensitivity to culturally diverse entities becomes increasingly im

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

ClinicalAgents: Multi-Agent Orchestration for Clinical Decision Making with Dual-Memory

DGX agent

arXiv:2603.26182v2 Announce Type: replace Abstract: While Large Language Models (LLMs) have demonstrated potential in healthcare, they often struggle with the complex, non-linear reasoning required fo

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

Cost-Sensitive Evaluation for Binary Classifiers

DGX agent

arXiv:2510.22016v2 Announce Type: replace Abstract: Selecting an appropriate evaluation metric for classifiers is crucial for model comparison, parameter optimization, and deployment decisions, yet th

model-releasesarxiv-cs-lg
28 May 2026
Model Releases

DeepSciVerify: Verifying Scientific Claim--Citation Alignment via LLM-Driven Evidence Escalation

DGX agent

arXiv:2605.27710v1 Announce Type: new Abstract: Misalignment between claims and their cited evidence is a common failure mode in reports generated by large language models, limiting their reliability

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Diffusion-Based Ukrainian Handwritten Text Generation with Cross-Domain Style Transfer

DGX agent

arXiv:2605.27487v1 Announce Type: cross Abstract: Handwritten text generation (HTG) conditioned on writer style has been widely studied for Latin scripts, but remains underexplored for low-resource an

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Do Audio LLMs Listen or Read? Analyzing and Mitigating Paralinguistic Failures with VoxParadox

DGX agent

arXiv:2605.27772v1 Announce Type: cross Abstract: Audio large language models (Audio LLMs) demonstrate strong performance on speech understanding tasks, yet their ability to understand paralinguistic

model-releasesarxiv-cs-lg
28 May 2026
Research

EAGer: Entropy-Aware GEneRation for Adaptive Inference-Time Scaling

DGX agent

arXiv:2510.11170v2 Announce Type: replace-cross Abstract: With the rise of reasoning language models and test-time scaling methods as a paradigm for improving model performance, substantial computatio

researcharxiv-cs-ai
28 May 2026
Model Releases

Efficient Pre-Training of LLMs through Truncated SVD Layers

DGX agent

arXiv:2605.28573v1 Announce Type: cross Abstract: The massive scaling of Large Language Models (LLMs) has made pretraining increasingly cost-prohibitive. While low-rank representation and orthonormal

model-releasesarxiv-cs-ai
28 May 2026
Safety

Escape the Language Prior: Mitigating Late-Stage Modality Collapse in Audio Reasoning via Modality-Aware Policy Optimization

DGX agent

arXiv:2605.27741v1 Announce Type: new Abstract: Audio and omni-modal large language models exhibit impressive cross-modal reasoning capabilities. However, applying standard reinforcement learning post

safetyarxiv-cs-cl
28 May 2026
Model Releases

Evolving Dataflow to process massive datasets for machine learning

DGX agent

Google created MapReduce more than 20 years ago to solve the scaling problems in data processing that the then young company was running into. The AI era that we are in now demands efficient, large-sc

model-releasesgoogle-cloud-ai
28 May 2026
Research

Explaining is Harder Than Predicting Alone: Evaluating Concept-based Explanations of MLLMs as ICL Visual Classifiers

DGX agent

arXiv:2605.28215v1 Announce Type: new Abstract: In-context learning (ICL) enables multimodal large language models (MLLMs) to classify images from a few labelled examples. Yet, how these models use th

researcharxiv-cs-ai
28 May 2026
Model Releases

HELEA: Hard-Negative Benchmark and LLM-based Reranking for Robust Entity Alignment

DGX agent

arXiv:2605.28308v1 Announce Type: new Abstract: Entity Alignment (EA) is essential for knowledge graph (KG) fusion, but existing benchmarks often allow models to exploit name overlap rather than relat

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

IAR2: Improving Autoregressive Visual Generation with Semantic-Detail Associated Token Prediction

DGX agent

arXiv:2510.06928v2 Announce Type: replace Abstract: Autoregressive models have emerged as a powerful paradigm for visual content creation, but often overlook the intrinsic structural properties of vis

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

IFMTBench: A Comprehensive Benchmark for Multilingual Translation Instruction Following

DGX agent

arXiv:2605.28218v1 Announce Type: new Abstract: Modern translation workflows demand more than semantic equivalence. Users routinely require models to preserve JSON or HTML schemas, honor curated gloss

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

Inpainting-Style Conditional Diffusion for Multivariable Time Series Forecasting

DGX agent

arXiv:2605.28324v1 Announce Type: new Abstract: In this paper, we propose a novel conditional diffusion-based framework for multivariable time-series solar power forecasting. The proposed method refor

model-releasesarxiv-cs-cv
28 May 2026
Agents

Knowing When to Ask: Segment-Level Credit Assignment for LLM Tool Use

DGX agent

arXiv:2605.27788v1 Announce Type: cross Abstract: Humans know when to reach for help e.g. 347 imes 28 warrants a calculator while 2+2 does not. Language models do not. Prompt-based approaches can inst

agentsarxiv-cs-cl
28 May 2026
Model Releases

Mags-RL: Wearing Multimodal LLMs a Magnifying Glass via Agentic Reinforcement Learning For Complex Scene Reasoning

DGX agent

arXiv:2605.27960v1 Announce Type: new Abstract: Despite their popularity and success, Multimodal Large Language Models (MLLMs) often struggle to interpret images accurately, which limits their reasoni

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

NL-MambaXCT: Self-Supervised Nested-Learning Mamba for Nomex Honeycomb X-ray CT Defect Classification

DGX agent

arXiv:2605.27454v1 Announce Type: cross Abstract: X-ray computed tomography (XCT) is widely used for non-destructive testing of Nomex honeycomb structures in aerospace manufacturing, but industrial in

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

Patched-DeltaNet: Token-Level Event-Driven Memory for Linear-Time Anomaly Detection

DGX agent

arXiv:2605.27992v1 Announce Type: new Abstract: Time series anomaly detection is critical for maintaining the reliability of mission-critical systems. While Transformer-based models like PatchTST have

model-releasesarxiv-cs-lg
28 May 2026
Tutorials

Random Process Flow Matching: Generative Implicit Representations of Multivariate Random Fields

DGX agent

arXiv:2605.28625v1 Announce Type: new Abstract: Generative modeling provides a powerful framework for learning data distributions. These models initially relied on probabilistic methods such as Gaussi

tutorialsarxiv-cs-lg
28 May 2026
Research

RULER: Representation-Level Verification of Machine Unlearning

DGX agent

arXiv:2605.27569v1 Announce Type: new Abstract: Machine unlearning aims to remove the influence of specific training records from a deployed model without retraining from scratch. Current protocols ve

researcharxiv-cs-ai
28 May 2026
Research

Satisfiability Solving with LLMs: A Matched-Pair Evaluation of Reasoning Capability

DGX agent

arXiv:2605.28602v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for tasks that implicitly reduce to Boolean satisfiability (SAT), yet their reasoning ability on SAT

researcharxiv-cs-ai
28 May 2026
Research

SelfJudge: Faster Speculative Decoding via Self-Supervised Judge Verification

DGX agent

arXiv:2510.02329v2 Announce Type: replace-cross Abstract: Speculative decoding accelerates LLM inference by verifying candidate tokens from a draft model against a larger target model. Recent judge de

researcharxiv-cs-ai
28 May 2026
Model Releases

Simorgh at SemEval-2026 task 7: Region-Aware Hybrid Retrieval for Low-Resource Cultural Reasoning in Multilingual Question Answering

DGX agent

arXiv:2605.27636v1 Announce Type: new Abstract: Although Large Language Models (LLMs) demonstrate excellent capabilities and performance for general reasoning tasks within the general public domain, t

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

SmartIterator: Visual Analytics Workflows for Supervising Unsupervised Data Grouping

DGX agent

arXiv:2605.28219v1 Announce Type: cross Abstract: Unsupervised learning methods -- topic modeling, partition-based and density-based clustering -- produce data groupings without human guidance, yet ch

model-releasesarxiv-cs-ai
28 May 2026
Local Ai

Soft Specialists: alpha-Renyi Ensembles for Uncertainty-Aware LLM Post-Training

DGX agent

arXiv:2605.27747v1 Announce Type: cross Abstract: Existing training approaches for large language models learn a single set of parameters, based on large volumes of data, which is typically heterogene

local-aiarxiv-cs-lg
28 May 2026
Research

Towards Reliable Multilingual LLMs-as-a-Judge: An Empirical Study

DGX agent

arXiv:2605.28710v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for the automatic evaluation of generated text, yet most prior work focuses on English. Despite the

researcharxiv-cs-ai
28 May 2026
Research

Universal Time Series Generation with Neural Controlled Differential Equations

DGX agent

arXiv:2605.28507v1 Announce Type: new Abstract: Recent work on the sequence universality of State Space Models (SSMs) has introduced efficient, maximally expressive continuous-time approaches for time

researcharxiv-cs-lg
28 May 2026
Model Releases

VULPO: Context-Aware Vulnerability Detection via On-Policy LLM Optimization

DGX agent

arXiv:2511.11896v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have recently shown strong potential in vulnerability detection (VD). However, accurately detecting vulnerabiliti

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

WeatherCity: Urban Scene Reconstruction with Controllable Multi-Weather Transformation

DGX agent

arXiv:2602.22096v2 Announce Type: replace Abstract: Editable high-fidelity 4D scenes are crucial for autonomous driving, as they can be applied to end-to-end training and closed-loop simulation. Howev

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

Axial-Centric Cross-Plane Attention for 3D Medical Image Classification

DGX agent

arXiv:2602.21636v2 Announce Type: replace Abstract: Abridged: Clinicians commonly interpret 3D medical images by examining multiple anatomical planes rather than relying on volumetric views. In clinic

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

Causal Representation Learning for Generalisable Recommendation

DGX agent

arXiv:2605.27043v1 Announce Type: cross Abstract: Predictive models trained on observational data often fail to generalise to the distributions they encounter when deployed, especially when the traini

model-releasesarxiv-cs-lg
27 May 2026
Model Releases

Drive-P2D: A Progressive Perception-to-Decision Benchmark for VLMs in Autonomous Driving

DGX agent

arXiv:2601.14702v2 Announce Type: replace Abstract: Autonomous driving requires reliable perception and safe decision-making in complex scenarios. Recent vision-language models (VLMs) demonstrate reas

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

ECHO-2: A Large-Scale Distributed Rollout Framework for Cost-Efficient Reinforcement Learning

DGX agent

arXiv:2602.02192v5 Announce Type: replace Abstract: Reinforcement learning (RL) is a critical stage in post-training large language models (LLMs), involving repeated interaction between rollout genera

model-releasesarxiv-cs-lg
27 May 2026
Model Releases

EpiCurveBench: Evaluating VLMs on Epidemic Curve Digitization

DGX agent

arXiv:2605.27195v1 Announce Type: new Abstract: Chart-to-data extraction with vision-language models (VLMs) is increasingly evaluated on benchmarks that show diminishing headroom (frontier VLMs exceed

model-releasesarxiv-cs-cl
27 May 2026
Research

Error Analysis of Discrete Flow with Generator Matching

DGX agent

arXiv:2509.21906v3 Announce Type: replace-cross Abstract: Discrete flow models offer a powerful framework for learning distributions over discrete state spaces and have demonstrated superior performan

researcharxiv-cs-lg
27 May 2026
Model Releases

InfoSynth: Information-Guided Benchmark Synthesis for LLMs

DGX agent

arXiv:2601.00575v2 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated significant advancements in reasoning and code generation, but efficiently creating new benchmarks to

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness

DGX agent

arXiv:2605.26438v1 Announce Type: cross Abstract: Large language models can recognize when they are being evaluated (evaluation awareness) and behave differently because of that, which undermines the

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Med-CoReasoner: Reducing Language Disparities in Medical Reasoning via Language-Informed Co-Reasoning

DGX agent

arXiv:2601.08267v3 Announce Type: replace Abstract: While reasoning-enhanced large language models perform strongly on English medical tasks, a persistent multilingual gap remains, with substantially

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

MobileMoE: Scaling On-Device Mixture of Experts

DGX agent

arXiv:2605.27358v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) has become the de facto architecture for hundred-billion-parameter language models, yet its advantages at sub-billion scales

model-releasesarxiv-cs-ai
27 May 2026
← Previous
1…455456457458459…1357
Next →