AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
29 Apr 2026

LegalMidm: Use-Case-Driven Legal Domain Specialization for Korean Large Language Model

ApplicationsDGX agent

arXiv:2604.25297v1 Announce Type: new Abstract: In recent years, the rapid proliferation of open-source large language models (LLMs) has spurred efforts to turn general-purpose models into domain spec

Less Is More: Fast and Accurate Reasoning with Cross-Head Unified Sparse Attention

ResearchDGX agent

arXiv:2508.07101v2 Announce Type: replace Abstract: Large reasoning models achieve strong performance through test-time scaling, but this incurs substantial computational overhead due to long decoding

Leverage Laws: A Per-Task Framework for Human-Agent Collaboration

AgentsDGX agent

arXiv:2604.25040v1 Announce Type: cross Abstract: We propose a per-task leverage ratio for human-agent collaboration: human work displaced by an agent, divided by the human time required to specify th


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Libra-VLA: Achieving Learning Equilibrium via Asynchronous Coarse-to-Fine Dual-System

SafetyDGX agent

arXiv:2604.24921v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models are a promising paradigm for generalist robotic manipulation by grounding high-level semantic instructions into ex

Limited Linguistic Diversity in Embodied AI Datasets

ResearchDGX agent

arXiv:2601.03136v2 Announce Type: replace Abstract: Language plays a critical role in Vision-Language-Action (VLA) models, yet the linguistic characteristics of the datasets used to train and evaluate

LLM-ReSum: A Framework for LLM Reflective Summarization through Self-Evaluation

Model ReleasesDGX agent

arXiv:2604.25665v1 Announce Type: new Abstract: Reliable evaluation of large language model (LLM)-generated summaries remains an open challenge, particularly across heterogeneous domains and document

LongSumEval: Question-Answering Based Evaluation and Feedback-Driven Refinement for Long Document Summarization

Model ReleasesDGX agent

arXiv:2604.25130v1 Announce Type: new Abstract: Evaluating long document summaries remains the primary bottleneck in summarization research. Existing metrics correlate weakly with human judgments and

Luminol-AIDetect: Fast Zero-shot Machine-Generated Text Detection based on Perplexity under Text Shuffling

Local AiDGX agent

arXiv:2604.25860v1 Announce Type: new Abstract: Machine-generated text (MGT) detection requires identifying structurally invariant signals across generation models, rather than relying on model-specif

MAIC-UI: Making Interactive Courseware with Generative UI

SafetyDGX agent

arXiv:2604.25806v1 Announce Type: new Abstract: Creating interactive STEM courseware traditionally requires HTML/CSS/JavaScript expertise, leaving barriers for educators. While generative AI can produ

Marco-MoE: Open Multilingual Mixture-of-Expert Language Models with Efficient Upcycling

ResearchDGX agent

arXiv:2604.25578v1 Announce Type: new Abstract: We present Marco-MoE, a suite of fully open multilingual sparse Mixture-of-Experts (MoE) models. Marco-MoE features a highly sparse design in which only

MGSM-Pro: A Simple Strategy for Robust Multilingual Mathematical Reasoning Evaluation

Model ReleasesDGX agent

arXiv:2601.21225v2 Announce Type: replace Abstract: Large language models have made substantial progress in mathematical reasoning. However, benchmark development for multilingual evaluation has lagge

MGTEVAL: An Interactive Platform for Systemtic Evaluation of Machine-Generated Text Detectors

ResearchDGX agent

arXiv:2604.25152v1 Announce Type: cross Abstract: We present MGTEVAL, an extensible platform for systematic evaluation of Machine-Generated Text (MGT) detectors. Despite rapid progress in MGT detectio

MiMo-Embodied: X-Embodied Foundation Model Technical Report

AgentsDGX agent

arXiv:2511.16518v2 Announce Type: replace-cross Abstract: We open-source MiMo-Embodied, the first cross-embodied foundation model to successfully integrate and achieve state-of-the-art performance in

Mitigating Coordinate Prediction Bias from Positional Encoding Failures

Model ReleasesDGX agent

arXiv:2510.22102v2 Announce Type: replace-cross Abstract: While Multimodal Large Language Models (MLLMs) excel at general vision-language tasks, precise coordinate prediction remains a significant cha

Modeling Human-Like Color Naming Behavior in Context

ResearchDGX agent

arXiv:2604.25674v1 Announce Type: new Abstract: Modeling the emergence of human-like lexicons in computational systems has advanced through the use of interacting neural agents, which simulate both le

MolReFlect: Towards In-Context Fine-grained Alignments between Molecules and Texts

SafetyDGX agent

arXiv:2411.14721v2 Announce Type: replace Abstract: Molecule discovery is a pivotal research field, impacting everything from medicine to materials. Recently, Large Language Models (LLMs) have been wi

Named Entity Recognition of Historical Texts via Large Language Model

ResearchDGX agent

arXiv:2508.18090v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have demonstrated remarkable versatility across a wide range of natural language processing tasks and domains. On

Navigating Global AI Regulation: A Multi-Jurisdictional Retrieval-Augmented Generation System

SafetyDGX agent

arXiv:2604.25448v1 Announce Type: new Abstract: Navigating AI regulation across jurisdictions is increasingly difficult for policymakers, legal professionals, and researchers. To address this, we pres

Odysseys: Benchmarking Web Agents on Realistic Long Horizon Tasks

Model ReleasesDGX agent

arXiv:2604.24964v1 Announce Type: cross Abstract: Existing web agent benchmarks have largely converged on short, single-site tasks that frontier models are approaching saturation on. However, real wor

OMHBench: Benchmarking Balanced and Grounded Omni-Modal Multi-Hop Reasoning

Model ReleasesDGX agent

arXiv:2508.16198v3 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have increasingly supported omni-modal processing across text, vision, and speech. However, existing evalua

One Refiner to Unlock Them All: Inference-Time Reasoning Elicitation via Reinforcement Query Refinement

SafetyDGX agent

arXiv:2604.25444v1 Announce Type: new Abstract: Large Language Models (LLMs) often fail to utilize their latent reasoning capabilities due to a distributional mismatch between ambiguous human inquirie

Phase-Associative Memory: Sequence Modeling in Complex Hilbert Space

Model ReleasesDGX agent

arXiv:2604.05030v2 Announce Type: replace Abstract: Experiments probing natural language processing by both humans and LLMs suggest that the meaning of a semantic expression is indeterminate prior to

PolyKV: A Shared Asymmetrically-Compressed KV Cache Pool for Multi-Agent LLM Inference

Model ReleasesDGX agent

arXiv:2604.24971v1 Announce Type: cross Abstract: We present PolyKV, a system in which multiple concurrent inference agents share a single, asymmetrically compressed KV cache pool. Rather than allocat

Praxy Voice: Voice-Prompt Recovery + BUPS for Commercial-Class Indic TTS from a Frozen Non-Indic Base at Zero Commercial-Training-Data Cost

Model ReleasesDGX agent

arXiv:2604.25441v1 Announce Type: cross Abstract: Commercial TTS systems produce near-native Indic audio, but the best open-source bases (Chatterbox, Indic Parler-TTS, IndicF5) trail them on measured

Principled Detection of Hallucinations in Large Language Models via Multiple Testing

ResearchDGX agent

arXiv:2508.18473v3 Announce Type: replace Abstract: While Large Language Models (LLMs) have emerged as powerful foundational models to solve a variety of tasks, they have also been shown to be prone t

Progressing beyond Art Masterpieces or Touristic Cliches: how to assess your LLMs for cultural alignment?

SafetyDGX agent

arXiv:2604.25654v1 Announce Type: new Abstract: Although the cultural (mis)alignment of Large Language Models (LLMs) has attracted increasing attention -- often framed in terms of cultural bias -- unt

PSI-Bench: Towards Clinically Grounded and Interpretable Evaluation of Depression Patient Simulators

Model ReleasesDGX agent

arXiv:2604.25840v1 Announce Type: new Abstract: Patient simulators are gaining traction in mental health training by providing scalable exposure to complex and sensitive patient interactions. Simulati

PSP: An Interpretable Per-Dimension Accent Benchmark for Indic Text-to-Speech

Model ReleasesDGX agent

arXiv:2604.25476v1 Announce Type: cross Abstract: Standard text-to-speech (TTS) evaluation measures intelligibility (WER, CER) and overall naturalness (MOS, UTMOS) but does not quantify accent. A synt

Quantifying and Mitigating Socially Desirable Responding in LLMs: A Desirability-Matched Graded Forced-Choice Psychometric Study

Model ReleasesDGX agent

arXiv:2602.17262v2 Announce Type: replace Abstract: Human self-report questionnaires are increasingly used in NLP to benchmark and audit large language models (LLMs), from persona consistency to safet

R^3-SQL: Ranking Reward and Resampling for Text-to-SQL

AgentsDGX agent

arXiv:2604.25325v1 Announce Type: cross Abstract: Modern Text-to-SQL systems generate multiple candidate SQL queries and rank them to judge a final prediction. However, existing methods face two limit

RbtAct: Rebuttal as Supervision for Actionable Review Feedback Generation

Model ReleasesDGX agent

arXiv:2603.09723v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used across the scientific workflow, including to draft peer-review reports. However, many AI-generate

Recursive Multi-Agent Systems

AgentsDGX agent

arXiv:2604.25917v1 Announce Type: cross Abstract: Recursive or looped language models have recently emerged as a new scaling axis by iteratively refining the same model computation over latent states

RELIC: Evaluating Complex Reasoning via the Recognition of Languages In-Context

ApplicationsDGX agent

arXiv:2506.05205v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used to solve complex tasks where they must retrieve and compose many pieces of in-context information

Rethinking Layer Redundancy in Large Language Models: Calibration Objectives and Search for Depth Pruning

ResearchDGX agent

arXiv:2604.24938v1 Announce Type: cross Abstract: Depth pruning improves the inference efficiency of large language models by removing Transformer blocks. Prior work has focused on importance criteria

Scaling Probabilistic Transformer via Efficient Cross-Scale Hyperparameter Transfer

Model ReleasesDGX agent

arXiv:2604.25409v1 Announce Type: new Abstract: Probabilistic Transformer (PT), a white-box probabilistic model for contextual word representation, has demonstrated substantial similarity to standard

SciDER: Scientific Data-centric End-to-end Researcher

AgentsDGX agent

arXiv:2603.01421v2 Announce Type: replace-cross Abstract: Automated scientific discovery with large language models is transforming the research lifecycle from ideation to experimentation, yet existin

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining

Model ReleasesDGX agent

arXiv:2602.10718v3 Announce Type: replace-cross Abstract: While FP8 attention has shown substantial promise in innovations like FlashAttention-3, its integration into the decoding phase of the DeepSee

Subliminal Steering: Stronger Encoding of Hidden Signals

SafetyDGX agent

arXiv:2604.25783v1 Announce Type: new Abstract: Subliminal learning describes a student language model inheriting a behavioral bias by fine-tuning on seemingly innocuous data generated by a biased tea

The Dynamics of Delusion: Modeling Bidirectional False Belief Amplification in Human-Chatbot Dialogue

ResearchDGX agent

arXiv:2604.25096v1 Announce Type: new Abstract: There is growing concern that AI chatbots might fuel delusional beliefs in users. Some have suggested that humans and chatbots mutually reinforce false

The Russian Legislative Corpus

SafetyDGX agent

arXiv:2406.04855v3 Announce Type: replace Abstract: We present a comprehensive corpus of Russian primary and secondary legislation adopted between 1991 and 2025, comprising 304,382 texts (194,425,905

The Structured Output Benchmark: A Multi-Source Benchmark for Evaluating Structured Output Quality in Large Language Models

Model ReleasesDGX agent

arXiv:2604.25359v1 Announce Type: new Abstract: Large Language Models are increasingly being deployed to extract structured data from unstructured and semi-structured sources: parsing invoices, medica

The Surprising Universality of LLM Outputs: A Real-Time Verification Primitive

Model ReleasesDGX agent

arXiv:2604.25634v1 Announce Type: cross Abstract: We report a striking statistical regularity in frontier LLM outputs that enables a CPU-only scoring primitive running at 2.6 microseconds per token, w

Thinking About Thinking: Evaluating Reasoning in Post-Trained Language Models

SafetyDGX agent

arXiv:2510.16340v2 Announce Type: replace Abstract: Recent advances in post-training techniques have endowed Large Language Models (LLMs) with enhanced capabilities for tackling complex, logic-intensi

Three Models of RLHF Annotation: Extension, Evidence, and Authority

SafetyDGX agent

arXiv:2604.25895v1 Announce Type: cross Abstract: Preference-based alignment methods, most prominently Reinforcement Learning with Human Feedback (RLHF), use the judgments of human annotators to shape

TouchAI: Exploring human-AI perceptual alignment in touch through language model representations

SafetyDGX agent

arXiv:2406.06587v2 Announce Type: replace Abstract: Aligning large language models (LLMs) behaviour with human intent is critical for future AI. An important yet often overlooked aspect of this alignm

Toward a Functional Geometric Algebra for Natural Language Semantics

ResearchDGX agent

arXiv:2604.25902v1 Announce Type: new Abstract: Distributional and neural approaches to natural language semantics have been built almost exclusively on conventional linear algebra: vectors, matrices,

Toward Multimodal Conversational AI for Age-Related Macular Degeneration

ResearchDGX agent

arXiv:2604.25720v1 Announce Type: cross Abstract: Despite strong performance of deep learning models in retinal disease detection, most systems produce static predictions without clinical reasoning or

Unrequited Emotions: Investigating the Gaps in Motivation and Practice in Speech Emotion Recognition Research

SafetyDGX agent

arXiv:2604.25776v1 Announce Type: new Abstract: Critical analyses of emotion recognition technology have raised ethical concerns around task validity and potential downstream impacts, urging researche

VLM Judges Can Rank but Cannot Score: Task-Dependent Uncertainty in Multimodal Evaluation

Model ReleasesDGX agent

arXiv:2604.25235v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly used as automated judges for multimodal systems, yet their scores provide no indication of reliability.

Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution

SafetyDGX agent

arXiv:2604.03472v2 Announce Type: replace Abstract: Co-evolutionary self-play, where one language model generates problems and another solves them, promises autonomous curriculum learning without huma

Voice, Bias, and Coreference: An Interpretability Study of Gender in Speech Translation

SafetyDGX agent

arXiv:2511.21517v2 Announce Type: replace Abstract: Unlike text, speech conveys information about the speaker, such as gender, through acoustic cues like pitch. This gives rise to modality-specific bi

VOYAGER: A Training Free Approach for Generating Diverse Datasets using LLMs

ResearchDGX agent

arXiv:2512.12072v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly being used to generate synthetic datasets for the evaluation and training of downstream models. Howeve

Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models

Model ReleasesDGX agent

arXiv:2604.25591v1 Announce Type: cross Abstract: Recent audio-aware large language models (ALLMs) have demonstrated strong capabilities across diverse audio understanding and reasoning tasks, but the

What Makes Good Instruction-Tuning Data? An In-Context Learning Perspective

ResearchDGX agent

arXiv:2604.25132v1 Announce Type: new Abstract: Instruction-tuning datasets often contain substantial redundancy and low-quality samples, necessitating effective data selection methods. We propose an

When Thoughts Meet Facts: Reusable Reasoning for Long-Context LMs

Model ReleasesDGX agent

arXiv:2510.07499v2 Announce Type: replace Abstract: Recent Long-Context Language Models (LCLMs) can process hundreds of thousands of tokens in a single prompt, enabling new opportunities for knowledge

WhisperPipe: A Resource-Efficient Streaming Architecture for Real-Time Automatic Speech Recognition

HardwareDGX agent

arXiv:2604.25611v1 Announce Type: new Abstract: Real-time automatic speech recognition (ASR) systems face a fundamental trade-off between transcription accuracy and computational efficiency, particula

Why Do LLM-based Web Agents Fail? A Hierarchical Planning Perspective

ResearchDGX agent

arXiv:2603.14248v2 Announce Type: replace-cross Abstract: Large language model (LLM) web agents are increasingly used for web navigation but remain far from human reliability on realistic, long-horizo

Why Does Reinforcement Learning Generalize? A Feature-Level Mechanistic Study of Post-Training in Large Language Models

ResearchDGX agent

arXiv:2604.25011v1 Announce Type: new Abstract: Reinforcement learning (RL)-based post-training often improves the reasoning performance of large language models (LLMs) beyond the training domain, whi

Wiki Dumps to Training Corpora: South Slavic Case

ResearchDGX agent

arXiv:2604.25384v1 Announce Type: new Abstract: This paper presents a methodology for transforming raw Wikimedia dumps into quality textual corpora for seven South Slavic languages. The work is divide

28 Apr 2026

A Benchmark Suite of Reddit-Derived Datasets for Mental Health Detection

Model ReleasesDGX agent

arXiv:2604.23458v1 Announce Type: new Abstract: The growing availability of online support groups has opened up new windows to study mental health through natural language processing (NLP). However, i

← Previous
1…9596979899…129
Next →