AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
26 May 2026

PolyGnosis 2.0: Enhancing LLM Reasoning via Agentic Harness Engineering for Polymarket and OSINT Insight Extraction

SafetyDGX agent

arXiv:2605.25958v1 Announce Type: new Abstract: This paper introduces PolyGnosis 2.0, a pioneering multi-agent architecture designed to extract predictive intelligence by synthesizing Polymarket anoma

PolySAE: Modeling Feature Interactions in Sparse Autoencoders via Polynomial Decoding

Model ReleasesDGX agent

arXiv:2602.01322v2 Announce Type: replace-cross Abstract: Sparse autoencoders (SAEs) interpret neural network representations by decomposing activations into sparse combinations of dictionary atoms. H

PowLU: An Activation Function for Stable Pre-Training of LLMs

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.25704v1 Announce Type: new Abstract: In contemporary large language models (LLMs), the swish-gated linear unit (SwiGLU) activation function is widely adopted to regulate the information flo

Prefix Teach, Suffix Fade: Local Teachability Collapse in Strong-to-Weak On-Policy Distillation

Local AiDGX agent

arXiv:2605.13643v2 Announce Type: replace Abstract: On-policy distillation (OPD) trains a student model on its own rollouts using dense feedback from a stronger teacher. Prior literature suggests that

Prism: A Plug-in Reproducible Infrastructure for Scalable Multimodal Continual Instruction Tuning

ApplicationsDGX agent

arXiv:2605.26110v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) achieve versatility by reformulating diverse tasks into a unified instruction-following framework via instruc

Proactive for Uncertainty: Cause-Aware Error Diagnosis and Interactive Clarification for Spoken Dialogue Systems

ResearchDGX agent

arXiv:2605.25404v1 Announce Type: new Abstract: Cascaded Automatic Speech Recognition -- Large Language Model (ASR-LLM) pipelines remain popular for industrial Spoken Dialogue Systems (SDS), primarily

Probability Distributions Computed by Autoregressive Transformers

ResearchDGX agent

arXiv:2510.27118v4 Announce Type: replace Abstract: Most expressivity results for transformers treat them as language recognizers -- devices that accept or reject strings -- rather than as they are us

Quantifying the Impact of Translation Errors on Multilingual LLM Evaluation

Model ReleasesDGX agent

arXiv:2605.24904v1 Announce Type: new Abstract: Machine-translated benchmarks are widely used to assess the multilingual capabilities of large language models (LLMs), yet translation errors in these b

QUEST: Training Frontier Deep Research Agents with Fully Synthetic Tasks

Model ReleasesDGX agent

arXiv:2605.24218v1 Announce Type: new Abstract: Deep research agents extend the role of search engines from retrieving keyword-matched pages to synthesizing knowledge, fundamentally changing how human

Re-defining Humor Data Objects for AI Humor Research

ResearchDGX agent

arXiv:2605.25171v1 Announce Type: new Abstract: In most existing AI humor research, humor was treated as either 'present' or 'not present.' We explore the concept of humor as a social interaction with

Reading, Not Thinking: Understanding and Bridging the Modality Gap When Text Becomes Pixels in Multimodal LLMs

SafetyDGX agent

arXiv:2603.09095v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) can process text presented as images, yet they often perform worse than when the same content is provided a

Reinforcement Learning from Denoising Feedback

SafetyDGX agent

arXiv:2605.25638v1 Announce Type: new Abstract: Policy loss estimation remains a fundamental and long-standing challenge in reinforcement learning (RL) for diffusion language models (dLLMs). We introd

Repeated Sequences Reveal Gaps between Large Language Models and Natural Language

ResearchDGX agent

arXiv:2605.24850v1 Announce Type: new Abstract: Evaluating whether large language models (LLMs) capture the structure of natural language beyond local fluency remains an open challenge. Existing evalu

Retrieval as Reasoning: Self-Evolving Agent-Native Retrieval via LLM-Wiki

AgentsDGX agent

arXiv:2605.25480v1 Announce Type: new Abstract: LLM agents require retrieval to behave less like one-shot context fetching and more like reasoning: searching, reading, traversing, and deciding when ev

Retrieved In-Context Principles from Previous Mistakes

ResearchDGX agent

arXiv:2407.05682v2 Announce Type: replace Abstract: In-context learning (ICL) has been instrumental in adapting Large Language Models (LLMs) to downstream tasks using correct input-output examples. Re

ROC Analysis for Evaluating Translation Quality Estimation Systems

ResearchDGX agent

arXiv:2605.24721v1 Announce Type: new Abstract: The increasing use of automated translation quality estimation (QE) systems calls for practical, decision-oriented methods for evaluating their performa

RotMoLE: Enhancing Mixture of Low-Rank Experts through Rotational Gating Mechanism

Model ReleasesDGX agent

arXiv:2605.25565v1 Announce Type: cross Abstract: While Large Language Models (LLMs) are commonly fine-tuned to handle domain-specific tasks before being applied to vertical applications, adapting the

RouteScan: A Non-Intrusive Approach to Auditing MoE LLMs Safety via Expert Routing Telemetry

SafetyDGX agent

arXiv:2605.24817v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) architectures have become an increasingly important paradigm for scaling Large Language Models (LLMs). As MoE models are incr

Rubato: Transcribing Piano Music with Timestamps

ResearchDGX agent

arXiv:2605.24291v1 Announce Type: cross Abstract: We consider the conversion of musical recordings into human-readable sheet music annotated with timestamps. Such output lets a listener clearly visual

Scaling Natural-Language Graph-Based Test Time Compute for Automated Theorem Proving

ResearchDGX agent

arXiv:2503.11657v3 Announce Type: replace Abstract: Large language models have demonstrated remarkable capabilities in natural language processing tasks requiring multi-step logical reasoning capabili

Schema-Grounded LLM Extraction for FHIR Patient Digital Twins

Model ReleasesDGX agent

arXiv:2601.05847v2 Announce Type: replace Abstract: We revisit the problem of constructing interoperable patient digital twins from unstructured electronic health records (EHRs) and argue that the tas

SEAL: Synergistic Co-Evolution of Agents and Learning Environments

SafetyDGX agent

arXiv:2605.24426v1 Announce Type: new Abstract: Large Language Model (LLM) agents are increasingly improved through interaction, yet most self-evolution methods adapt either the policy or the learning

Selective Latent Thinking: Adaptive Compression of LLM Reasoning Chains

SafetyDGX agent

arXiv:2605.25745v1 Announce Type: new Abstract: Explicit chain-of-thought (CoT) reasoning substantially improves the reasoning ability of large language models (LLMs), but incurs high inference cost d

SLAP: Stratified Loss-based Pruning for On-Policy Data-Efficient Instruction Tuning

Model ReleasesDGX agent

arXiv:2605.23969v1 Announce Type: new Abstract: Instruction tuning has optimized the specialized capabilities of large language models (LLMs), but it often requires extensive datasets and prolonged tr

SliceWorld: A Predictive and Controllable World-State Model for CT Report Generation

SafetyDGX agent

arXiv:2605.24371v1 Announce Type: cross Abstract: CT report generation (CTRG) requires models to summarize three-dimensional anatomical context and pathological findings from hundreds of axial slices.

Spiking the training data to correct for test set contamination

ResearchDGX agent

arXiv:2605.24818v1 Announce Type: cross Abstract: The literature on test set contamination largely focuses on detection, but the correction of contaminated test scores is underexplored. Our core propo

StepGap: A Hybrid NLI-LLM Checker for Step-Level Evidence-Gap Detectionin Multi-Hop Question Answering

ResearchDGX agent

arXiv:2605.24733v1 Announce Type: new Abstract: We present extbf{StepGap}, a hybrid NLI-LLM decision tree that detects step-level evidence gaps in multi-hop QA and emits one of three typed labels: ext

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models

AgentsDGX agent

arXiv:2605.26014v1 Announce Type: cross Abstract: Many video reasoning tasks require tracking motion, temporal order, and evolving visual states across frames. Existing methods built on large vision-l

StreamProfileBench: A Benchmark for Fine-Grained User Profile Inference in Real-World Streaming Scenarios

Model ReleasesDGX agent

arXiv:2605.25758v1 Announce Type: new Abstract: Large Language Models (LLMs) have reshaped user profiling, yet current evaluations mainly focus on static data snapshots. This paradigm overlooks the re

Structural Abstraction as an Inductive Bias for Non-Stationary Language Model Training

Model ReleasesDGX agent

arXiv:2603.17198v2 Announce Type: replace-cross Abstract: A foundational principle in cognitive science holds that intelligent agents do not learn by storing experiences as isolated instances, but by

Structure-Aware RAG: Structured Retrieval Augmented Generation from Noisy Data for Conversational Agents

ApplicationsDGX agent

arXiv:2605.24366v1 Announce Type: new Abstract: Large Language Models (LLMs) have been widely adopted in conversational applications. However, their reliance on parametric knowledge limits reliability

SURGE: On the Potential of Large Language Models as General-Purpose Surrogate Code Executors

Model ReleasesDGX agent

arXiv:2502.11167v5 Announce Type: replace-cross Abstract: Neural surrogate models are powerful and efficient tools in data mining. Meanwhile, large language models (LLMs) have demonstrated remarkable

Testing the Deliteralization Hypothesis in Human and Machine Translation

ResearchDGX agent

arXiv:2605.25686v1 Announce Type: new Abstract: The recent shift from dedicated NMT systems to general-purpose LLMs has reshaped machine translation, with LLMs reported to produce more fluent, less li

Thaka at KSAA-2026 Task 2: Regularized Fine-Tuning for Arabic Speech Diacritization

ResearchDGX agent

arXiv:2605.25928v1 Announce Type: new Abstract: We describe the winning system for Task 2 of the KSAA-2026 Shared Task on Arabic Speech Dictation with Automatic Diacritization. The task requires produ

The Age of Curiosity Meets the Age of AI: Benchmarking Child Safety in Large Language Models

Model ReleasesDGX agent

arXiv:2605.25510v1 Announce Type: new Abstract: Children increasingly have access to Large Language Models (LLMs), which may expose them to responses that are developmentally inappropriate or require

The LSCD Benchmark: a Testbed for Diachronic Word Meaning Tasks

Model ReleasesDGX agent

arXiv:2404.00176v3 Announce Type: replace Abstract: Lexical Semantic Change Detection (LSCD) is a complex, lemma-level task, which is usually operationalized based on two subsequently applied usage-le

The meaning of prompts and the prompts of meaning: Semiotic reflections and modelling

ResearchDGX agent

arXiv:2509.14250v2 Announce Type: replace Abstract: This paper explores prompts and prompting in large language models (LLMs) as dynamic semiotic phenomena, drawing on Peirce's triadic model of signs,

The Multilingual Curse at the Retrieval Layer: Evidence from Amharic

ResearchDGX agent

arXiv:2605.24556v1 Announce Type: cross Abstract: Multilingual retrieval increasingly underpins cross-lingual question answering and retrieval-augmented generation. Strong zero-shot scores on multilin

The Tokenizer Tax Across 25 European Languages: Domain Invariance, Cross-Lingual Few-Shot Effects, and the Ukrainian Penalty

ResearchDGX agent

arXiv:2605.24718v1 Announce Type: new Abstract: Tokenizer fertility the number of tokens per word imposes a hidden cost on non-English NLP. We measure fertility for ten foundation models across 25 Eur

They Are Not the Same: Direct Causes Are Not Grounded Emotion Explanations

ResearchDGX agent

arXiv:2605.25208v1 Announce Type: new Abstract: Emotion-Cause Pair Extraction (ECPE) was introduced to explain why an emotion occurs, but this goal is now often reduced to binary pair/non-pair predict

TimeSpot: Benchmarking Geo-Temporal Understanding in Vision-Language Models in Real-World Settings

Model ReleasesDGX agent

arXiv:2603.06687v2 Announce Type: replace-cross Abstract: Geo-temporal understanding, the ability to infer location, time, and contextual properties from visual input alone, underpins applications suc

Tool-Call Dependency Structure is Linearly Decodable in LLM Agent Residual Streams

AgentsDGX agent

arXiv:2605.25310v1 Announce Type: new Abstract: Tool-using LLM agents produce trajectories whose calls form a directed dependency graph: earlier tool outputs supply arguments to later calls. Whether t

Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content

SafetyDGX agent

arXiv:2509.12672v2 Announce Type: replace Abstract: The volume of machine-generated content online has grown dramatically due to the widespread use of Large Language Models (LLMs), leading to new chal

Toxicity in Twitch Chats: An LLM-Based Analysis Across Gaming Communities

ResearchDGX agent

arXiv:2605.24000v1 Announce Type: new Abstract: Toxicity in online gaming communities remains a persistent challenge, manifesting across genres, platforms, and player interactions. While much research

TRACE: A taxonomy-grounded synthetic dataset for teaching-program generation and session interpretation in Applied Behavior Analysis

Model ReleasesDGX agent

arXiv:2605.25038v1 Announce Type: new Abstract: Applied Behavior Analysis (ABA) is a clinical discipline whose documentation, teaching programs and multi-session behavioral logs, is formulaic and high

Trait-Aware Policy Optimization for Autoregressive Multi-Trait Essay Scoring

SafetyDGX agent

arXiv:2605.25731v1 Announce Type: new Abstract: Multi-trait essay scoring aims to provide fine-grained evaluation of writing quality across multiple dimensions. However, how to effectively post-train

Transformers over-extend what humans underlearn: the case of Spanish L-shaped morphome

ApplicationsDGX agent

arXiv:2507.21556v3 Announce Type: replace Abstract: The cognitive reality of irregular morphological patterns has been debated for decades: do speakers extend them to novel forms, or are they lexical

Translators as Invisible Teachers of AI: Copyright, Translation Memory, and the Political Economy of Linguistic Data

ApplicationsDGX agent

arXiv:2605.24842v1 Announce Type: new Abstract: This paper examines how the labour of translators has been transformed into foundational data capital for the age of artificial intelligence (AI). Trans

Triplet-Block Diffusion RWKV

ResearchDGX agent

arXiv:2605.25969v1 Announce Type: new Abstract: Causal Transformer language models suffer from strictly sequential decoding and a quadratic per-step attention cost. While linear-time causal models and

TypedCSIP: Typed Counterfactual Pretraining for Chinese Legislative Conflict Classification

Model ReleasesDGX agent

arXiv:2605.25474v1 Announce Type: new Abstract: TypedCSIP is a typed counterfactual pretraining method for the conflict-classification task of the LCR-CN benchmark (Zhao et al., 2026): given a (superi

Universal Activation Verbalizer: A Unified Framework for Cross-Model Activation Explanation

ResearchDGX agent

arXiv:2605.25903v1 Announce Type: new Abstract: Activation verbalization explains hidden representations in natural language, but existing methods are mostly limited to self-explanation, where each mo

Universal Boosts, Specific Suppressors: Sparse Autoencoder Steering of Medical Vision-Language Models

SafetyDGX agent

arXiv:2605.24977v1 Announce Type: cross Abstract: Medical vision-language models (VLMs) often hallucinate findings when generating chest X-ray reports: they fabricate findings that are not present in

Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval

ApplicationsDGX agent

arXiv:2605.24530v1 Announce Type: new Abstract: Document retrieval in real-world scenarios faces significant challenges due to diverse document formats and modalities. Traditional text-based approache

What Are We Actually Decoding? Source Attribution for Non-Invasive Brain-to-Language Retrieval

Local AiDGX agent

arXiv:2605.24524v1 Announce Type: cross Abstract: In non-invasive neural language decoding, results can be inflated by sources that are not stimulus-evoked neural evidence: decoder priors, embedding-b

What Makes a Medical Checker Trainable? Diagnosing Signal Collapse and Reward Hacking in Checker-Guided RAG for Biomedical QA

Model ReleasesDGX agent

arXiv:2605.25988v1 Announce Type: new Abstract: Medical RAG needs evidence-grounded claims, so plugging a claim-level NLI checker into retrieval-augmented RL is intuitive. extbf{We find that the check

What Questions Should Robots Be Able to Answer? A Dataset of User Questions for Explainable Robotics

ResearchDGX agent

arXiv:2510.16435v2 Announce Type: replace-cross Abstract: With the growing use of large language models and conversational interfaces in human-robot interaction, robots' ability to answer user questio

When Do LLM Agents Treat Surface Noise Differently from Semantic Noise? A 68-Cell Measurement Study with a Held-Out Trace-Level Validation

Model ReleasesDGX agent

arXiv:2605.25981v1 Announce Type: new Abstract: We document an empirical phenomenon in chain-of-thought and ReAct agents driven by ten large language models from seven architecture families: meaning-b

When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift

ResearchDGX agent

arXiv:2605.25629v1 Announce Type: new Abstract: Weak-to-strong (W2S) generalization is a promising framework for scalable oversight, yet existing evaluations often test students under matched train--t

When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards

SafetyDGX agent

arXiv:2605.25864v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved remarkable advancements in reasoning capabilities empowered by Reinforcement Learning with Verifiable Rewar

WhenLoss: Diagnosing Write and Retrieval Bottlenecks in Long-Context Memory Systems

Model ReleasesDGX agent

arXiv:2605.24579v1 Announce Type: new Abstract: Long-context memory systems often fail under fixed budgets, but end-to-end evaluation does not reveal whether evidence was discarded during compression

← Previous
1…6162636465…129
Next →