AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
5 Jun 2026

Staged Factorial Screening for Budget-Constrained Micro-Pretraining

HardwareDGX agent

arXiv:2606.05186v1 Announce Type: cross Abstract: Budget-constrained micro-pretraining often requires triaging many candidate recipes on a shared accelerator before larger search budgets are spent. We

Statistical Priors for Implicit Preferences: Decoupling Skill Selection as a Local Harness in Personal Agents

Local AiDGX agent

arXiv:2606.05828v1 Announce Type: cross Abstract: As Large Language Model (LLM) capabilities advance, locally deployed personal agents relying on API-based remote models and external skills have emerg

Statistically Reliable LLM-Based Ranking Evaluation via Prediction-Powered Inference

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.05308v1 Announce Type: cross Abstract: With PRECISE, we extended Prediction-Powered Inference to produce bias-corrected estimates of ranking evaluation metrics by combining a small human-la

Staying with the Uncertainty: Uncertainty-Scaffolding Strategies for Artificial Moral Advisors in LLM-to-LLM Simulated Conversations

AgentsDGX agent

arXiv:2606.05890v1 Announce Type: new Abstract: LLMs are increasingly deployed as Artificial Moral Advisors (AMA) in a variety of contexts: what kind of conversational patterns should they display? In

SubtleMemory: A Benchmark for Fine-Grained Relational Memory Discrimination in Long-Horizon AI Agents

Model ReleasesDGX agent

arXiv:2606.05761v1 Announce Type: cross Abstract: Persistent AI assistants, such as OpenClaw, accumulate large collections of related memories over long-term interactions. As these memories grow, they

TARPO: Token-Wise Latent-Explicit Reasoning via Action-Routing Policy Optimization

Model ReleasesDGX agent

arXiv:2606.05859v1 Announce Type: new Abstract: Latent reasoning has emerged as a promising alternative to discrete Chain-of-Thought (CoT) in large language models (LLMs), enabling more expressive rea

Temporal Preference Concepts and their Functions in a Large Language Model

Local AiDGX agent

arXiv:2606.05194v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly being deployed to make decisions that require trading off near-term gains against long-term consequences

Ten Headache Specialists versus Artificial Intelligence for Clinical Literature Summarization: A Critical Evaluation and Comparison

Model ReleasesDGX agent

arXiv:2606.05436v1 Announce Type: cross Abstract: Summarizing the latest medical literature to guide clinical decision-making is essential for evidence-based medicine and high-quality patient care. Ye

TensorBench: Benchmarking Coding Agents on a Compiler-Based Tensor Framework

Model ReleasesDGX agent

arXiv:2606.05570v1 Announce Type: new Abstract: Repository-level coding benchmarks face a trade-off between task difficulty and evaluation reliability: tasks that challenge frontier models often invol

The Generator-Eraser Paradox: Community Guidelines for Responsible LLM-Assisted Dialect Resource Creation

ApplicationsDGX agent

arXiv:2606.06004v1 Announce Type: new Abstract: Dialect resources occupy a unique position at the intersection of scientific description, cultural preservation, and computational infrastructure. Large

The Granularity Gap: A Multi-Dimensional Longitudinal Audit of Sycophancy in Gemini Models

Model ReleasesDGX agent

arXiv:2606.05183v1 Announce Type: new Abstract: Large language models are increasingly deployed as high-stakes advisors, yet standard alignment benchmarks treat sycophancy as a binary failure mode. We

The Mirage of Performance Gains: Why Contrastive Decoding Fails to Mitigate Object Hallucinations in MLLMs?

Model ReleasesDGX agent

arXiv:2504.10020v4 Announce Type: replace Abstract: Contrastive decoding strategies are widely used to reduce object hallucinations in multimodal large language models (MLLMs). These methods work by c

The Prosody of Emojis

ApplicationsDGX agent

arXiv:2508.00537v2 Announce Type: replace Abstract: Prosodic features such as pitch, timing, and intonation are central to spoken communication, conveying emotion, intent, and discourse structure. In

The Self-Correction Illusion: LLMs Correct Others but Not Themselves

AgentsDGX agent

arXiv:2606.05976v1 Announce Type: cross Abstract: Recent work shows that LLM agents struggle to correct errors in their own reasoning traces yet show markedly higher correction rates when identical cl

The Tell-Tale Norm: ell_2 Magnitude as a Signal for Reasoning Dynamics in Large Language Models

ResearchDGX agent

arXiv:2606.06188v1 Announce Type: new Abstract: Recent work has sought to understand Large Language Models (LLMs) reasoning, yet a principled, model-intrinsic signal that captures its layer-wise reaso

To Be Multimodal or Not to Be: Query-Adaptive Audio-Visual Person Retrieval via Active Modality Detection

ApplicationsDGX agent

arXiv:2606.05931v1 Announce Type: new Abstract: When retrieving a person from a video archive by voice and face, should the system be multimodal or not? In real-world broadcast archives, unlike curate

Toward Culturally Aligned LLMs through Ontology-Guided Multi-Agent Reasoning

SafetyDGX agent

arXiv:2601.21700v3 Announce Type: replace Abstract: Large Language Models (LLMs) increasingly support culturally sensitive decision making, yet often exhibit misalignment due to skewed pretraining dat

Towards Truly Multilingual ASR: Generalizing Code-Switching ASR to Unseen Language Pairs

ResearchDGX agent

arXiv:2606.05846v1 Announce Type: new Abstract: Automatic Speech Recognition (ASR) has become a key technology for human--AI interaction. However, code-switching ASR (CS-ASR) remains particularly chal

Trajectory Dynamics in Language Model Hidden States Predict Human Processing Costs Beyond Surprisal

Local AiDGX agent

arXiv:2606.05346v1 Announce Type: new Abstract: Human language comprehension unfolds sequentially: each word is processed in the context of those that came before, and the interpretation builds increm

UNIVID: Unified Vision-Language Model for Video Moderation

SafetyDGX agent

arXiv:2606.05748v1 Announce Type: cross Abstract: Global-scale video moderation faces a dual challenge: the need for fine-grained multi-modal reasoning and the demand for interpretable outputs to supp

Unsupervised Skill Discovery for Agentic Data Analysis

AgentsDGX agent

arXiv:2606.06416v1 Announce Type: cross Abstract: Inference-time skill augmentation provides a lightweight way to improve data-analytic agents by injecting reusable procedural knowledge without updati

USAD 2.0: Scaling Representation Distillation for Universal Audio Understanding

ResearchDGX agent

arXiv:2606.06444v1 Announce Type: cross Abstract: Audio encoders are critical to modern audio applications as large language models (LLMs) increasingly rely on a single encoder for diverse inputs. Whi

Using Large Language Models to Support High Volume Application Review for an Undergraduate Research Program

Model ReleasesDGX agent

arXiv:2606.05564v1 Announce Type: new Abstract: Undergraduate research programs such as the Summer Undergraduate Research Fellowship (SURF) at Purdue University receive thousands of applications every

Value-and-Structure Alignment for Routing-Consistent Quantization of Mixture-of-Experts Models

SafetyDGX agent

arXiv:2606.05688v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models scale foundation models efficiently by activating only a subset of experts for each token, but their large number of exp

Vavanagi: a Community-run Platform for Documentation of the Hula Language in Papua New Guinea

ResearchDGX agent

arXiv:2603.14210v2 Announce Type: replace Abstract: We present Vavanagi, a community-run platform for Hula (Vula'a), an Austronesian language of Papua New Guinea with approximately 10,000 speakers. Va

What Makes Two Language Models Think Alike?

ResearchDGX agent

arXiv:2406.12620v3 Announce Type: replace Abstract: Do architectural and training differences influence the way models represent and process language? Traditional similarity metrics tell us whether tw

What's in a Name? Morphological Shortcuts by LLMs in Pharmacology

SafetyDGX agent

arXiv:2606.05616v1 Announce Type: new Abstract: The morphological form of a word can often give cues to its meaning, but purely relying on these mappings can lead to overgeneralization in high-stakes

When AI Says It Feels

SafetyDGX agent

arXiv:2606.05734v1 Announce Type: cross Abstract: Large language models (LLMs) are generally constrained from expressing feelings through human-preference alignment in post-training processes. This po

When Evidence is Sparse: Weakly Supervised Early Failure Alerting in Dialogs and LLM-Agent Trajectories

SafetyDGX agent

arXiv:2606.05414v1 Announce Type: new Abstract: Early failure alerting requires deciding, while a dialog or agent trajectory is still unfolding, whether to flag it as likely to fail. This is challengi

When New Generators Arrive: Lifelong Machine-Generated Text Attribution via Ridge Feature Transfer

ResearchDGX agent

arXiv:2606.05626v1 Announce Type: new Abstract: Machine-generated text (MGT) attribution aims to identify the specific generator responsible for a given text, thereby providing fine-grained evidence f

Where does Absolute Position come from in decoder-only Transformers?

ResearchDGX agent

arXiv:2606.06160v1 Announce Type: cross Abstract: RoPE-trained transformers distinguish absolute position in their attention patterns, even though RoPE encodes only relative offsets in the inner produ

Would you still call this Dax? Novel Visual References in VLMs and Humans

Model ReleasesDGX agent

arXiv:2606.05409v1 Announce Type: cross Abstract: Vision-language models (VLMs), like human learners, are frequently exposed to new visual concepts, but how they map novel visual references to languag

You Only Index Once: Cross-Layer Sparse Attention with Shared Routing

ResearchDGX agent

arXiv:2606.06467v1 Announce Type: new Abstract: Long-context inference in modern LLMs is increasingly constrained by decoding efficiency, especially in reasoning-heavy settings where models generate l

YouZhi: Towards High-Concurrency Financial LLMs via Adaptive GQA-to-MLA Transition

Model ReleasesDGX agent

arXiv:2606.05868v1 Announce Type: new Abstract: Large language models (LLMs) drive significant financial innovations, yet their high-concurrency deployment is severely bottlenecked by KV cache memory

4 Jun 2026

A French Corpus Annotated for Multiword Expressions with Adverbial Function

ResearchDGX agent

arXiv:2606.04828v1 Announce Type: new Abstract: This paper presents a French corpus annotated for multiword expressions (MWEs) with adverbial function. This corpus is designed for investigation on inf

A Systematic Evaluation of Positional Bias in Multi-Video Summarization with MLLMs

Model ReleasesDGX agent

arXiv:2606.04596v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) are increasingly used for video understanding, yet their reliability under multi-video inputs remains poorly un

ACAT: A Collaborative Platform for Efficient Aspect-Based Sentiment Dataset Annotation

ResearchDGX agent

arXiv:2606.04189v1 Announce Type: new Abstract: Aspect-Based Sentiment Analysis (ABSA) requires high-quality datasets to train reliable models. However, existing annotation tools treat output as flat

Activation-Based Active Learning for In-Context Learning: Challenges and Insights

Model ReleasesDGX agent

arXiv:2606.05134v1 Announce Type: new Abstract: Deep active learning has previously been explored for LLM in-context sample selection, but not with methods that utilise recent advances in understandin

Adaptive Information Control for Search-Augmented LLM Reasoning

SafetyDGX agent

arXiv:2602.01672v2 Announce Type: replace Abstract: Search-augmented reasoning agents interleave multi-step reasoning with external retrieval, but uncontrolled retrieval can introduce redundant eviden

Agent Planning Benchmark: A Diagnostic Framework for Planning Capabilities in LLM Agents

Model ReleasesDGX agent

arXiv:2606.04874v1 Announce Type: new Abstract: Planning is central to LLM agents: before acting, an agent must decompose goals, select tools, reason over constraints, and decide when a task is infeas

AUDDT: A Unified Benchmark Toolkit for Audio and Speech Deepfake Detectors

Model ReleasesDGX agent

arXiv:2509.21597v2 Announce Type: replace-cross Abstract: With the prevalence of artificial intelligence (AI)-generated content, such as audio deepfakes, a large body of recent work has focused on dev

Automated Lexical Coverage for Language Learning: From General to Specialized Word Lists

ResearchDGX agent

arXiv:2512.15552v2 Announce Type: replace Abstract: A General Service List (GSL) is a commonly used resource for language learners to identify important English words. Traditional GSL creation is reso

BEATS: Bootstrapping E-commerce Attribute Taxonomies for Search through Iterative Human-AI Collaboration

ApplicationsDGX agent

arXiv:2606.04909v1 Announce Type: cross Abstract: E-commerce platforms in emerging markets often operate with underdeveloped product catalogs that contain only category taxonomies but lack structured

Benchmarking Living-Screen-Native GUI Agents on Short-Video Platforms

Model ReleasesDGX agent

arXiv:2606.04701v1 Announce Type: cross Abstract: GUI agents today assume a static screen, where the world is frozen between two actions. However, real interfaces such as short-video applications viol

Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation

AgentsDGX agent

arXiv:2510.13272v3 Announce Type: replace Abstract: Inspired by the success of reinforcement learning (RL) in Large Language Model (LLM) training for domains like math and code, recent work has begun

Beyond Retrieval: Learning Compact User Representations for Scalable LLM Personalization

Model ReleasesDGX agent

arXiv:2606.04547v1 Announce Type: cross Abstract: Personalizing large language models requires adapting model behavior to individual users while preserving robustness and deployment-scale efficiency.

Beyond Text Following: Repairable Arbitration Reversals in Audio-Language Models

ResearchDGX agent

arXiv:2606.05161v1 Announce Type: cross Abstract: Audio-language models (ALMs) often follow text that conflicts with audio, even when the audio evidence is clear. This raises a basic question: is the

Boosting Self-Consistency with Ranking

ResearchDGX agent

arXiv:2606.05054v1 Announce Type: new Abstract: Self-consistency improves large language models by sampling multiple reasoning paths and selecting the most frequent answer, but majority voting often f

BreastGPT: A Multimodal Large Language Model for the Full Spectrum of Breast Cancer Clinical Routine

Model ReleasesDGX agent

arXiv:2606.04911v1 Announce Type: cross Abstract: Breast cancer remains a leading cause of cancer-related mortality among women. Its clinical management requires multimodal reasoning across a clinical

Caliper: Probing Lexical Anchors versus Causal Structure in LLMs

Model ReleasesDGX agent

arXiv:2606.04915v1 Announce Type: new Abstract: Large language models reach 50 to 70% accuracy on causal reasoning benchmarks such as CLadder, but it is unclear whether this reflects structural reason

Can Crowdsourcing Survive the LLM Era? A Community Survey on Human Data Collection

TutorialsDGX agent

arXiv:2606.04924v1 Announce Type: new Abstract: The widespread use of Large Language Models (LLMs) as writing tools challenges the validity of crowdsourced data, as crowdworkers may outsource tasks to

Can Large Language Models Generalize Procedures Across Representations?

Model ReleasesDGX agent

arXiv:2602.03542v2 Announce Type: replace Abstract: Large language models (LLMs) are trained and tested extensively on symbolic representations such as code and graphs, yet real-world user tasks are o

Cartridges at Scale: Training Modular KV Caches over Large Document Collections

HardwareDGX agent

arXiv:2606.04557v1 Announce Type: new Abstract: Large Language Models can reason over long contexts, yet prefilling millions of tokens is wasteful as much of the content remains static across queries.

Characterizing, Evaluating, and Optimizing Complex Reasoning

TutorialsDGX agent

arXiv:2602.08498v2 Announce Type: replace Abstract: Large Reasoning Models (LRMs) increasingly rely on reasoning traces with complex internal structures. However, existing work lacks a unified answer

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding

ResearchDGX agent

arXiv:2606.04418v1 Announce Type: cross Abstract: Neural audio codecs are a key component of speech processing pipelines, compressing audio into discrete tokens for downstream modeling. However, exist

Clinical Assistant for Remote Engagement Link (CARE-link): A Web-Based Electronic Health Records Software for Managing Diabetes

ResearchDGX agent

arXiv:2606.04952v1 Announce Type: cross Abstract: CARE-link is an open-source, web-based clinical support platform designed to improve the management of gestational diabetes by linking clinicians and

Computational conceptual history of scientific concepts: From early digital methods to LLMs

ResearchDGX agent

arXiv:2606.04118v1 Announce Type: new Abstract: This article situates large language models (LLMs) within the longer history of computational approaches to concept analysis in the history, philosophy,

Confidence Before Answering: A Paradigm Shift for Efficient LLM Uncertainty Estimation

SafetyDGX agent

arXiv:2603.05881v2 Announce Type: replace Abstract: Reliable deployment of large language models (LLMs) requires accurate uncertainty estimation. Existing methods are predominantly answer-first, produ

Covert Influence Between Language Models

SafetyDGX agent

arXiv:2606.04071v1 Announce Type: cross Abstract: As language models increasingly consume one another's outputs, covert influence -- a phenomenon where a sender's payload (the behavioral disposition i

CRAFT: Cost-aware Refinement And Front-aware Tuning of Prompts

ResearchDGX agent

arXiv:2606.04661v1 Announce Type: new Abstract: Prompts tuned for accuracy often grow long, raising inference cost on every model call. The best accuracy-cost trade-off depends on the task and the bud

← Previous
1…4344454647…129
Next →