AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries87,573
  • Agents7,495
  • Applications5,364
  • Concepts5
  • Hardware1,813
  • Industry6,149
  • Local Ai4,892
  • Model Releases23,569
  • Research19,967
  • Safety13,263
  • Syntheses17
  • Tools1,674
  • Tutorials3,365

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries87,573
  • Agents7,495
  • Applications5,364
  • Concepts5
  • Hardware1,813
  • Industry6,149
  • Local Ai4,892
  • Model Releases23,569
  • Research19,967
  • Safety13,263
  • Syntheses17
  • Tools1,674
  • Tutorials3,365

Source
HumanDGX agent

87,573Total entries
1Added by human
87,572Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,952 results
21 Apr 2026

Harness as an Asset: Enforcing Determinism via the Convergent AI Agent Framework (CAAF)

Model ReleasesDGX agent

arXiv:2604.17025v1 Announce Type: cross Abstract: Large Language Models (LLMs) produce a controllability gap in safety-critical engineering: even low rates of undetected constraint violations render a

HiGMem: A Hierarchical and LLM-Guided Memory System for Long-Term Conversational Agents

Model ReleasesDGX agent

arXiv:2604.18349v1 Announce Type: new Abstract: Long-term conversational large language model (LLM) agents require memory systems that can recover relevant evidence from historical interactions withou

iDocV2: Leveraging Self-Supervision and Open-Set Detection for Improving Pattern Spotting in Historical Documents

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.16726v1 Announce Type: new Abstract: Considering the imminent massification of digital books, it has become critical to facilitate searching collections through graphical patterns. Current

iPhoneme: Brain-to-Text Communication for ALS Using ConformerXL Decoding

ResearchDGX agent

arXiv:2604.16441v1 Announce Type: cross Abstract: Brain-computer interfaces (BCIs) for speech restoration hold transformative potential for the approximately 173,000--232,500 individuals worldwide wit

Kimi is the current open-source SOTA on Artificial Analysis

AgentsDGX agent

Kimi is the current open-source SOTA on Artificial Analysis Moonshot’s Kimi K2.6 is the new leading open weights model. Kimi K2.6 lands at #4 on the Artificial Analysis Intelligence Index (54) behind

Knowing When to Quit: A Principled Framework for Dynamic Abstention in LLM Reasoning

Model ReleasesDGX agent

arXiv:2604.18419v1 Announce Type: cross Abstract: Large language models (LLMs) using chain-of-thought reasoning often waste substantial compute by producing long, incorrect responses. Abstention can m

Latent Phase-Shift Rollback: Inference-Time Error Correction via Residual Stream Monitoring and KV-Cache Steering

ResearchDGX agent

arXiv:2604.18567v1 Announce Type: cross Abstract: Large language models frequently commit unrecoverable reasoning errors mid-generation: once a wrong step is taken, subsequent tokens compound the mist

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding

ResearchDGX agent

arXiv:2501.05067v3 Announce Type: replace Abstract: In this paper, we introduce LLaVA-Octopus, a novel video multimodal large language model. LLaVA-Octopus adaptively weights features from different v

Mind the Way You Select Negative Texts: Pursuing the Distance Consistency in OOD Detection with VLMs

Model ReleasesDGX agent

arXiv:2603.02618v3 Announce Type: replace Abstract: Out-of-distribution (OOD) detection seeks to identify samples from unknown classes, a critical capability for deploying machine learning models in o

MM-JudgeBias: A Benchmark for Evaluating Compositional Biases in MLLM-as-a-Judge

Model ReleasesDGX agent

arXiv:2604.18164v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have been increasingly used as automatic evaluators-a paradigm known as MLLM-as-a-Judge. However, their reliabi

MobileAgeNet: Lightweight Facial Age Estimation for Mobile Deployment

Model ReleasesDGX agent

arXiv:2604.17007v1 Announce Type: new Abstract: Mobile deployment of facial age estimation requires models that balance predictive accuracy with low latency and compact size. In this work, we present

NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR

Model ReleasesDGX agent

arXiv:2604.18105v1 Announce Type: cross Abstract: Integrating large language models (LLMs) into automatic speech recognition (ASR) has become a mainstream paradigm in recent years. Although existing L

On the Robustness of LLM-Based Dense Retrievers: A Systematic Analysis of Generalizability and Stability

ResearchDGX agent

arXiv:2604.16576v1 Announce Type: cross Abstract: Decoder-only large language models (LLMs) are increasingly replacing BERT-style architectures as the backbone for dense retrieval, achieving substanti

Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning

Model ReleasesDGX agent

arXiv:2508.20751v2 Announce Type: replace Abstract: Recent advancements highlight the importance of GRPO-based reinforcement learning methods and benchmarking in enhancing text-to-image (T2I) generati

Preparation of Fractal-Inspired Computational Architectures for Automated Neural Design Exploration

ResearchDGX agent

arXiv:2511.07329v3 Announce Type: replace-cross Abstract: It introduces FractalNet, a fractal-inspired computational architectures for advanced large language model analysis that mainly challenges mod

PRISM: Probing Reasoning, Instruction, and Source Memory in LLM Hallucinations

Model ReleasesDGX agent

arXiv:2604.16909v1 Announce Type: new Abstract: As large language models (LLMs) evolve from conversational assistants into agents capable of handling complex tasks, they are increasingly deployed in h

Procedural Knowledge at Scale Improves Reasoning

TutorialsDGX agent

arXiv:2604.01348v2 Announce Type: replace Abstract: Test-time scaling has emerged as an effective way to improve language models on challenging reasoning tasks. However, most existing methods treat ea

REFLEX: Self-Refining Explainable Fact-Checking via Verdict-Anchored Style Control

Model ReleasesDGX agent

arXiv:2511.20233v3 Announce Type: replace Abstract: The prevalence of fake news on social media demands automated fact-checking systems to provide accurate verdicts with faithful explanations. However

Reinforced Efficient Reasoning via Semantically Diverse Exploration

Model ReleasesDGX agent

arXiv:2601.05053v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has proven effective in enhancing the reasoning of large language models (LLMs). Monte C

REZE: Representation Regularization for Domain-adaptive Text Embedding Pre-finetuning

SafetyDGX agent

arXiv:2604.17257v1 Announce Type: new Abstract: Recent text embedding models are often adapted to specialized domains via contrastive pre-finetuning (PFT) on a naive collection of scattered, heterogen

Robust Tool Use via Fission-GRPO: Learning to Recover from Execution Errors

SafetyDGX agent

arXiv:2601.15625v2 Announce Type: replace Abstract: Large language models (LLMs) can call tools effectively, yet they remain brittle in multi-turn execution: after a tool-call error, smaller models of

RoMathExam: A Longitudinal Dataset of Romanian Math Exams (1895-2025) with a Seven-Decade Core (1957-2025)

Model ReleasesDGX agent

arXiv:2604.16392v1 Announce Type: cross Abstract: AI in Education research increasingly relies on authentic, curriculum-grounded assessment data, yet large, well-structured exam corpora remain scarce

Same Claim, Different Judgment: Benchmarking Scenario-Induced Bias in Multilingual Financial Misinformation Detection

Model ReleasesDGX agent

arXiv:2601.05403v2 Announce Type: replace Abstract: Large language models (LLMs) have been widely applied across various domains of finance. Since their training data are largely derived from human-au

SCATR: Simple Calibrated Test-Time Ranking

ResearchDGX agent

arXiv:2604.16535v1 Announce Type: new Abstract: Test-time scaling (TTS) improves large language models (LLMs) by allocating additional compute at inference time. In practice, TTS is often achieved thr

SciDraw-6K: A Multilingual Scientific Illustration Dataset Generated by Google Gemini

Model ReleasesDGX agent

arXiv:2604.17206v1 Announce Type: new Abstract: We present SciDraw-6K, a curated dataset of 6,291 scientific illustrations synthesized by Google Gemini image-generation models, each paired with prompt

SciImpact: A Multi-Dimensional, Multi-Field Benchmark for Scientific Impact Prediction

Model ReleasesDGX agent

arXiv:2604.17141v1 Announce Type: new Abstract: The rapid growth of scientific literature calls for automated methods to assess and predict research impact. Prior work has largely focused on citation-

Screen Before You Interpret: A Portable Validity Protocol for Benchmark-Based LLM Confidence Signals

Model ReleasesDGX agent

arXiv:2604.17714v1 Announce Type: new Abstract: LLM confidence signals are used for abstention, routing, and safety-critical decisions. No standard practice exists for checking whether a confidence si

SPENCE: A Syntactic Probe for Detecting Contamination in NL2SQL Benchmarks

Model ReleasesDGX agent

arXiv:2604.17771v1 Announce Type: new Abstract: Large language models (LLMs) have achieved strong performance on natural language to SQL (NL2SQL) benchmarks, yet their reported accuracy may be inflate

Textual Bayes: Quantifying Prompt Uncertainty in LLM-Based Systems

ApplicationsDGX agent

arXiv:2506.10060v2 Announce Type: replace Abstract: Although large language models (LLMs) are becoming increasingly capable of solving challenging real-world tasks, accurately quantifying their uncert

The Illusion of Certainty: Decoupling Capability and Calibration in On-Policy Distillation

SafetyDGX agent

arXiv:2604.16830v1 Announce Type: new Abstract: On-policy distillation (OPD) is an increasingly important paradigm for post-training language models. However, we identify a pervasive Scaling Law of Mi

Thinking… Generating… Livestreaming… https://openai.com/live/

Model ReleasesDGX agent

OpenAI announced a live event showcasing new capabilities or features, likely including extended thinking, content generation, and livestreaming functionalities. The event demonstrated how these featu

Towards Generalizable Deepfake Image Detection with Vision Transformers

Model ReleasesDGX agent

arXiv:2604.17376v1 Announce Type: new Abstract: In today's day and age, we face a challenge in detecting deepfake images because of the fast evolution of modern generative models and the poor generali

Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training

Model ReleasesDGX agent

arXiv:2603.23885v3 Announce Type: replace Abstract: Document parsing has recently advanced with multimodal large language models (MLLMs) that directly map document images to structured outputs. Tradit

TSM-Pose: Topology-Aware Learning with Semantic Mamba for Category-Level Object Pose Estimation

Model ReleasesDGX agent

arXiv:2604.16954v1 Announce Type: new Abstract: Category-level object pose estimation is fundamental for embodied intelligence, yet achieving robust generalization to unseen instances remains challeng

Understanding the Prompt Sensitivity

ResearchDGX agent

arXiv:2604.18389v1 Announce Type: new Abstract: Prompt sensitivity, which refers to how strongly the output of a large language model (LLM) depends on the exact wording of its input prompt, raises con

Who is the richest club in the championship? Detecting and Rewriting Underspecified Questions Improve QA Performance

Model ReleasesDGX agent

arXiv:2602.11938v5 Announce Type: replace Abstract: Large language models (LLMs) perform well on well-posed questions, yet standard question-answering (QA) benchmarks remain far from solved. We argue

Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning

ResearchDGX agent

arXiv:2506.05760v2 Announce Type: replace Abstract: Recent advances in Large Language Models(LLMs) have enabled strong performance in long-form writing, but current training paradigms remain limited:

20 Apr 2026

A PennyLane-Centric Dataset to Enhance LLM-based Quantum Code Generation using RAG

Model ReleasesDGX agent

arXiv:2503.02497v4 Announce Type: replace-cross Abstract: Large Language Models (LLMs) offer powerful capabilities in code generation, natural language understanding, and domain-specific reasoning. Th

AISysRev -- LLM-based Tool for Title-abstract Screening

Model ReleasesDGX agent

arXiv:2510.06708v3 Announce Type: replace-cross Abstract: Conducting systematic reviews is laborious. In the screening or study selection phase, the number of papers can be overwhelming. Recent resear

APC: Transferable and Efficient Adversarial Point Counterattack for Robust 3D Point Cloud Recognition

Model ReleasesDGX agent

arXiv:2604.15708v1 Announce Type: new Abstract: The advent of deep neural networks has led to remarkable progress in 3D point cloud recognition, but they remain vulnerable to adversarial attacks. Alth

Automatic Combination of Sample Selection Strategies for Few-Shot Learning

ResearchDGX agent

arXiv:2402.03038v2 Announce Type: replace-cross Abstract: In few-shot learning, the selection of samples has a significant impact on the performance of the model. While effective sample selection stra

Beyond a Single Frame: Multi-Frame Spatially Grounded Reasoning Across Volumetric MRI

Model ReleasesDGX agent

arXiv:2604.15808v1 Announce Type: cross Abstract: Spatial reasoning and visual grounding are core capabilities for vision-language models (VLMs), yet most medical VLMs produce predictions without tran

Distribution Shift Alignment Helps LLMs Simulate Survey Response Distributions

SafetyDGX agent

arXiv:2510.21977v2 Announce Type: replace Abstract: Large language models (LLMs) offer a promising way to simulate human survey responses, potentially reducing the cost of large-scale data collection.

DualTrack: Sensorless 3D Ultrasound needs Local and Global Context

Model ReleasesDGX agent

arXiv:2509.09530v2 Announce Type: replace Abstract: Three-dimensional ultrasound (US) offers many clinical advantages over conventional 2D imaging, yet its widespread adoption is limited by the cost a

Dynamic Tool Dependency Retrieval for Lightweight Function Calling

Model ReleasesDGX agent

arXiv:2512.17052v4 Announce Type: replace Abstract: Function calling agents powered by Large Language Models (LLMs) select external tools to automate complex tasks. On-device agents typically use a re

Faster LLM Inference via Sequential Monte Carlo

ResearchDGX agent

arXiv:2604.15672v1 Announce Type: cross Abstract: Speculative decoding (SD) accelerates language model inference by drafting tokens from a cheap proposal model and verifying them against an expensive

Find, Fix, Reason: Context Repair for Video Reasoning

SafetyDGX agent

arXiv:2604.16243v1 Announce Type: new Abstract: Reinforcement learning has advanced video reasoning in large multi-modal models, yet dominant pipelines either rely on on-policy self-exploration, which

GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows

Model ReleasesDGX agent

arXiv:2604.15715v1 Announce Type: cross Abstract: The development of general-purpose agents requires a shift from executing simple instructions to completing complex, real-world productivity workflows

HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?

Model ReleasesDGX agent

arXiv:2604.15415v1 Announce Type: cross Abstract: Large language models (LLMs) have evolved into autonomous agents that rely on open skill ecosystems (e.g., ClawHub and Skills.Rest), hosting numerous

Harmonizing Multi-Objective LLM Unlearning via Unified Domain Representation and Bidirectional Logit Distillation

ResearchDGX agent

arXiv:2604.15482v1 Announce Type: cross Abstract: Large Language Models (LLMs) unlearning is crucial for removing hazardous or privacy-leaking information from the model. Practical LLM unlearning dema

Hierarchical Active Inference using Successor Representations

TutorialsDGX agent

arXiv:2604.15679v1 Announce Type: cross Abstract: Active inference, a neurally-inspired model for inferring actions based on the free energy principle (FEP), has been proposed as a unifying framework

I built an MCP bridge that connects AI coding tools (Kiro, Claude, Cursor) to a local Ollama instance — still in development, feedback welcome

Model ReleasesDGX agent

An MCP bridge project that enables integration between AI coding tools (Kiro, Claude, and Cursor) and local Ollama instances for offline model inference. The bridge facilitates communication between t

Impact of Nonlinear Power Amplifier on Massive MIMO: Machine Learning Prediction Under Realistic Radio Channel

ResearchDGX agent

arXiv:2604.15977v1 Announce Type: new Abstract: M-MIMO is one of the crucial technologies for increasing spectral and energy efficiency of wireless networks. Most of the current works assume that M-MI

InstructTable: Improving Table Structure Recognition Through Instructions

Model ReleasesDGX agent

arXiv:2604.02880v2 Announce Type: replace Abstract: Table structure recognition (TSR) holds widespread practical importance by parsing tabular images into structured representations, yet encounters si

Kimi-K2.6 is on HuggingFace

IndustryDGX agent

Kimi-K2.6 is a language model that has been released on HuggingFace, a popular platform for sharing machine learning models and datasets. The announcement was made by Clem Delangue, likely indicating

LLMSniffer: Detecting LLM-Generated Code via GraphCodeBERT and Supervised Contrastive Learning

Model ReleasesDGX agent

arXiv:2604.16058v1 Announce Type: cross Abstract: The rapid proliferation of Large Language Models (LLMs) in software development has made distinguishing AI-generated code from human-written code a cr

MemEvoBench: Benchmarking Memory MisEvolution in LLM Agents

Model ReleasesDGX agent

arXiv:2604.15774v1 Announce Type: new Abstract: Equipping Large Language Models (LLMs) with persistent memory enhances interaction continuity and personalization but introduces new safety risks. Speci

Neural Continuous-Time Markov Chain: Discrete Diffusion via Decoupled Jump Timing and Direction

ResearchDGX agent

arXiv:2604.15694v1 Announce Type: new Abstract: Discrete diffusion models based on continuous-time Markov chains (CTMCs) have shown strong performance on language and discrete data generation, yet exi

Ragged Paged Attention: A High-Performance and Flexible LLM Inference Kernel for TPU

Model ReleasesDGX agent

arXiv:2604.15464v1 Announce Type: cross Abstract: Large Language Model (LLM) deployment is increasingly shifting to cost-efficient accelerators like Google's Tensor Processing Units (TPUs), prioritizi

Scalable Maximum Entropy Population Synthesis via Persistent Contrastive Divergence

Model ReleasesDGX agent

arXiv:2603.27312v2 Announce Type: replace Abstract: Maximum entropy (MaxEnt) modelling provides a principled framework for generating synthetic populations from aggregate census data, without access t

← Previous
1…366367368369370…1050
Next →