AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
4 Jun 2026

Rethinking Continual Experience Internalization for Self-Evolving LLM Agents

SafetyDGX agent

arXiv:2606.04703v1 Announce Type: new Abstract: Experience internalization converts contextual experience from past interactions into reusable parametric capability, offering a promising path toward c

SAID: Accelerating Diffusion-Based Language Models via Scaffold-Aware Iterative Decoding

ResearchDGX agent

arXiv:2606.04974v1 Announce Type: new Abstract: Diffusion large language models (DLLMs) enable non-autoregressive generation by iteratively denoising corrupted token sequences with bidirectional conte

SANE Schema-aware Natural-language Evaluation of Biological Data

ResearchDGX agent

arXiv:2606.04500v1 Announce Type: new Abstract: High-throughput microscopy generates large, structured datasets capturing cellular responses to pharmacological perturbations, but accessing these datas


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Segment, Embed, and Align: A Universal Recipe for Aligning Subtitles to Signing

SafetyDGX agent

arXiv:2512.08094v2 Announce Type: replace Abstract: The goal of this work is to develop a universal approach for aligning subtitles (i.e., spoken language text with corresponding timestamps) to contin

Self-Evaluation Is Already There: Eliciting Latent Judge Calibration in Base LLMs with Minimal Data

Local AiDGX agent

arXiv:2606.05122v1 Announce Type: new Abstract: Large language models are increasingly evaluated by other models, raising a natural question: can a model predict how a judge will score its own output?

SemBlock: Semantic Boundary Dynamic Blocks for Diffusion LLMs

Local AiDGX agent

arXiv:2606.04964v1 Announce Type: new Abstract: Diffusion language models (DLMs) generate text through iterative denoising, and blockwise decoding improves their practicality by committing tokens in l

SMADE-IE: Sparse Multi-Agent Framework with Evidence-Driven Debate for Zero-Shot Information Extraction

Model ReleasesDGX agent

arXiv:2606.04691v1 Announce Type: new Abstract: Zero-shot information extraction (IE) with large language models (LLMs) has attracted increasing attention due to its flexibility in adapting to new sch

SocialCoach: Personalized Social Skill Learning with RL-based Agentic Tutoring and Practice

SafetyDGX agent

arXiv:2606.04155v1 Announce Type: cross Abstract: Social skills such as negotiation and leadership are crucial for personal and professional success in today's interconnected world. However, scalable

Solving Zebra Puzzles Using Constraint-Guided Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2407.03956v3 Announce Type: replace-cross Abstract: Prior research has enhanced the ability of Large Language Models (LLMs) to solve logic puzzles using techniques such as chain-of-thought promp

SparDA: Sparse Decoupled Attention for Efficient Long-Context LLM Inference

HardwareDGX agent

arXiv:2606.04511v1 Announce Type: new Abstract: Sparse attention reduces compute and memory bandwidth for long-context LLM inference. However, two key challenges remain: (1) KV cache capacity still gr

SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space

SafetyDGX agent

arXiv:2511.20102v3 Announce Type: replace Abstract: Sparse attention reduces the quadratic complexity of full self-attention but faces two challenges: (1) an attention gap, where applying sparse atten

Stateful Visual Encoders for Vision-Language Models

AgentsDGX agent

arXiv:2606.04433v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly used in multi-image, multi-turn agentic settings where decisions depend on visual changes. However, in

Stepwise Reasoning Enhancement for LLMs via External Subgraph Generation

Model ReleasesDGX agent

arXiv:2606.04454v1 Announce Type: new Abstract: Large language models have shown strong performance in natural language generation and downstream reasoning tasks, but they still struggle with logical

STRIDE: Training Data Attribution via Sparse Recovery from Subset Perturbations

Model ReleasesDGX agent

arXiv:2606.05165v1 Announce Type: cross Abstract: Training Data Attribution (TDA) seeks to trace a model's predictions back to its training data. The gold standard for TDA relies on causal interventio

TaDA: Calibrated Probe Gating for Task-Domain LoRA Merging

Model ReleasesDGX agent

arXiv:2606.05016v1 Announce Type: new Abstract: Combining a task LoRA adapter with a domain LoRA adapter into a single unified model is a practical yet largely unexplored challenge. Existing methods t

The Mechanistic Emergence of Symbol Grounding in Language Models

ApplicationsDGX agent

arXiv:2510.13796v3 Announce Type: replace Abstract: Symbol grounding (Harnad, 1990) describes how symbols such as words acquire their meanings by connecting to real-world sensorimotor experiences. Rec

Training-Free Lexical-Dense Fusion for Conversational-Memory Retrieval

ResearchDGX agent

arXiv:2606.04194v1 Announce Type: cross Abstract: Retrieving the few past turns that answer a new query across long multi-session histories is the retrieval bottleneck behind long-term conversational

Translation Heads: Disentangling meaning from language in LLM-based machine translation

ResearchDGX agent

arXiv:2602.04613v2 Announce Type: replace Abstract: Mechanistic Interpretability (MI) seeks to explain how neural networks implement their capabilities, but the scale of Large Language Models (LLMs) h

T^star: Progressive Block Scaling for Masked Diffusion Language Models Through Trajectory Aware Reinforcement Learning

ResearchDGX agent

arXiv:2601.11214v5 Announce Type: replace Abstract: We present T^star, a simple TraceRL-based training curriculum for progressive block-size scaling in masked diffusion language models (MDMs). Startin

UniFine: A Unified and Fine-grained Approach for Zero-shot Vision-Language Understanding

ResearchDGX agent

arXiv:2307.00862v3 Announce Type: replace-cross Abstract: Vision-language tasks, such as VQA, SNLI-VE, and VCR are challenging because they require the model's reasoning ability to understand the sema

Using Text-Based Causal Inference to Disentangle Factors Influencing Online Review Ratings

ApplicationsDGX agent

arXiv:2606.04286v1 Announce Type: new Abstract: Online reviews provide valuable insights into the perceived quality of facets of a product or service. While aspect-based sentiment analysis has focused

Validity Threats for Foundation Model Research

ResearchDGX agent

arXiv:2606.05029v1 Announce Type: cross Abstract: Controlled experiments are the backbone of machine learning research, but at the scale of modern foundation models, they have become prohibitively exp

VCIFBench: Evaluating Complex Instruction Following for Video Understanding

Model ReleasesDGX agent

arXiv:2606.04588v1 Announce Type: new Abstract: Multimodal large language models have made rapid progress in video understanding, yet existing benchmarks largely rely on simple prompts and provide lim

VentAgent: When LLMs Learn to Breathe -- Multi-Objective Arbitration for ARDS Ventilation

SafetyDGX agent

arXiv:2606.04632v1 Announce Type: cross Abstract: Mechanical ventilation for Acute Respiratory Distress Syndrome (ARDS) requires balancing competing physiological goals, including oxygenation, lung pr

Video2LoRA: Parametric Video Internalization for Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.04351v1 Announce Type: cross Abstract: Processing video in vision-language models is expensive: each frame occupies hundreds of tokens, and inference cost scales with every frame and every

WETBench: A Benchmark for Detecting Task-Specific Machine-Generated Text on Wikipedia

Model ReleasesDGX agent

arXiv:2507.03373v2 Announce Type: replace Abstract: Given Wikipedia's role as a trusted source of high-quality, reliable content, concerns are growing about the proliferation of low-quality machine-ge

When Clients Stop Following: A Cognitive Conceptualization Diagram-driven Framework for Strategic Counseling

ResearchDGX agent

arXiv:2606.04389v1 Announce Type: new Abstract: Large Language Models (LLMs) show promise in psychological counseling, yet existing benchmarks rely heavily on highly cooperative simulated clients. We

When Retrieval Doesn't Help: A Large-Scale Study of Biomedical RAG

ResearchDGX agent

arXiv:2606.04127v1 Announce Type: new Abstract: Medical question answering is a high-stakes setting where factual errors can have serious consequences. Retrieval-augmented generation (RAG) is widely v

3 Jun 2026

A cross-domain tropical species dataset with Chinese vernacular names and CITES source links

ResearchDGX agent

arXiv:2606.03156v1 Announce Type: new Abstract: We describe a versioned cross-domain dataset of 410,499 active tropical species (working snapshot 2026-04-20) spanning three applied subdomains -- tropi

A Locally Deployed RAG-Based Academic Advising System for Course Selection

TutorialsDGX agent

arXiv:2606.02983v1 Announce Type: new Abstract: The correct sequence of courses in the curriculum based on prerequisites between courses is of great importance for students to develop their knowledge

A Pocket Offline Model for Simultaneous Speech Translation as CUNI Submission to IWSLT 2026

SafetyDGX agent

arXiv:2606.03948v1 Announce Type: new Abstract: We implement simultaneous translation capability with the offline direct speech-to-text translation model Canary, using the state-of-the-art policy Alig

ARBOR: Online Process Rewards via a Reusable Rubric Buffer for Search Agents

Local AiDGX agent

arXiv:2606.03239v1 Announce Type: new Abstract: LLM-based search agents are trained predominantly with outcome-only reward, leaving the search process itself unsupervised. This signal degenerates on o

Assessing Pause Thresholds for empirical Translation Process Research

ApplicationsDGX agent

arXiv:2604.01410v2 Announce Type: replace Abstract: Text production (and translations) proceeds in the form of stretches of typing, interrupted by keystroke pauses. It is often assumed that fast typin

AutoTail-BSFGM: Class-Balance-Aware Fine-Tuning for Chinese Scholarly Text Classification

ResearchDGX agent

arXiv:2606.03576v1 Announce Type: new Abstract: Scholarly text classification supports literature organization, subject indexing, and research intelligence, but Chinese scholarly corpora often contain

Backdoor Unlearning Generalization: A Path Toward the Removal of Unknown Triggers in LLMs

SafetyDGX agent

arXiv:2606.03785v1 Announce Type: new Abstract: Backdoor attacks in Large Language Models (LLMs) are a growing security concern, where models can generate adversary-chosen content. Existing defenses t

Benchmarking Speech-to-Speech Translation Models

ApplicationsDGX agent

arXiv:2606.03241v1 Announce Type: new Abstract: Speech-to-speech translation (S2ST) has advanced rapidly, but offline evaluation lacks a unified protocol: studies report non-overlapping metric subsets

Beyond Ideal Instruction: A Comprehensive Framework for Evaluating LLMs in Realistic Interactions

Model ReleasesDGX agent

arXiv:2606.03318v1 Announce Type: new Abstract: Despite great advances in tool-use capabilities of large language models (LLMs), existing evaluation benchmarks struggle to fully align with real-world

Beyond Semantics: Modeling Factual and Affective Perceptual Experiences from Vision-Language Data

ResearchDGX agent

arXiv:2606.03345v1 Announce Type: cross Abstract: We present P-Topics (Perception Topics) modeling, a novel problem for understanding how images are perceived affectively and across cultures. The goal

Beyond the Literal: Decomposing Pragmatic Intent in Multimodal Meme Understanding

ResearchDGX agent

arXiv:2606.03604v1 Announce Type: new Abstract: When asked what a meme or sarcastic post means, Large Vision Language Models (LVLMs) tend to describe what the image shows rather than what the author i

Beyond 'To whom it may concern': Tailoring Machine Translation to Audience and Intent

ResearchDGX agent

arXiv:2606.03259v1 Announce Type: new Abstract: Translation quality depends on purpose: the same source text demands different translations depending on audience, tone, and communicative intent. Yet M

Breaking the Self-Confirming Loop: Diagnosing and Mitigating Systemic Reward Bias in Self-Rewarding RL

SafetyDGX agent

arXiv:2510.08977v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) efficiently scales the reasoning ability of large language models (LLMs) but is bottlene

Can Factual Opinions Be Edited (Manipulated) in Large Language Models?

Model ReleasesDGX agent

arXiv:2606.03096v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly integrated into various domains, making knowledge editing techniques crucial yet potentially hazardous. Cu

Can LLM Rerankers Predict Their Own Ranking Performance?

ResearchDGX agent

arXiv:2606.03535v1 Announce Type: cross Abstract: Retrieval effectiveness varies substantially across queries, making it important to estimate ranking quality before relevance judgments are available.

Can Structural Cues Save LLMs? Evaluating Language Models in Massive Document Streams

Model ReleasesDGX agent

arXiv:2603.19250v2 Announce Type: replace Abstract: Evaluating language models in streaming environments is critical, yet underexplored. Existing benchmarks either focus on single complex events or pr

CAPER: Clause-Aligned Process Supervision for Text-to-SQL

Model ReleasesDGX agent

arXiv:2606.03327v1 Announce Type: cross Abstract: Text-to-SQL systems are typically evaluated by query-level execution correctness, but this terminal signal provides little guidance about which interm

Chatbots Output Meaningful (but Problematic) Language

Model ReleasesDGX agent

arXiv:2606.02973v1 Announce Type: new Abstract: Are utterances by AI chatbots meaningful? Concretely, if a user asks, say, Anthropic's agent Claude, 'What is the capital of Spain?' and Claude answers,

Coherence Maximization Improves Pluralistic Alignment

SafetyDGX agent

arXiv:2606.03110v1 Announce Type: new Abstract: Aligning AI systems with diverse human values requires value specifications grounded in concrete examples, but generating such examples without extensiv

Core-based Hierarchies for Efficient GraphRAG

ApplicationsDGX agent

arXiv:2603.05207v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) enhances large language models by incorporating external knowledge. However, existing vector-based method

CourseTimeQA: A Lecture-Video Benchmark and a Latency-Constrained Cross-Modal Fusion Method for Timestamped QA

Model ReleasesDGX agent

arXiv:2512.00360v2 Announce Type: replace Abstract: We study timestamped question answering over educational lecture videos under a single-GPU latency/memory budget. Given a natural-language query, th

CTR-Sink: Attention Sink for Language Models in Click-Through Rate Prediction

ResearchDGX agent

arXiv:2508.03668v2 Announce Type: replace Abstract: Click-Through Rate (CTR) prediction, a core task in recommendation systems, estimates user click likelihood using historical behavioral data. Modeli

DMT-CBT: Longitudinal Therapeutic State Modeling for CBT Counseling

Local AiDGX agent

arXiv:2606.03132v1 Announce Type: new Abstract: Large language models (LLMs) have shown growing potential for Cognitive Behavioral Therapy (CBT) counseling. However, most existing approaches still for

Do Value Vectors in Deep Layers Need Context from the Residual Stream?

Model ReleasesDGX agent

arXiv:2606.02780v1 Announce Type: new Abstract: The success of the transformer architecture as the backbone of modern LLMs is in large part due to its use of attention layers. An attention layer follo

Does Language Shift Break Medical Vision-Language Models? Indonesian Radiology Visual Question Answering Case Study

ApplicationsDGX agent

arXiv:2606.03693v1 Announce Type: new Abstract: Medical Vision-Language Models (VLMs) are typically evaluated on English radiology visual question answering benchmarks, leaving their robustness under

Don't Forget Your Embeddings: Robust Knowledge Erasure via Precise Editing of Embeddings

Model ReleasesDGX agent

arXiv:2606.03695v1 Announce Type: new Abstract: As language models are increasingly deployed in real-world applications, the ability to erase specific knowledge from them becomes critical for safety a

Dynamic Short Convolutions Improve Transformers

SafetyDGX agent

arXiv:2606.03825v1 Announce Type: cross Abstract: Transformers have become the dominant architecture for large language models, largely due to the scalability and flexibility of attention, feed-forwar

Enhancing Paraphrase Type Generation: The Impact of DPO and RLHF Evaluated with Human-Ranked Data

ResearchDGX agent

arXiv:2506.02018v2 Announce Type: replace Abstract: Paraphrasing re-expresses meaning to enhance applications like text simplification, machine translation, and question-answering. Specific paraphrase

Entropy Gate: Entropy Quenching for Near-Lossless Token Compression in LLM Pipelines

AgentsDGX agent

arXiv:2606.03739v1 Announce Type: new Abstract: LLM pipelines waste substantial token budgets on low-information content: repeated context, verbose responses, and redundant boilerplate. We introduce E

EntSQL: A Benchmark for Grounding Text-to-SQL in Long-Context Enterprise Knowledge

Model ReleasesDGX agent

arXiv:2606.03363v1 Announce Type: new Abstract: Text-to-SQL enables natural language access to databases, and recent LLMs have substantially advanced its capabilities. Existing benchmarks such as Spid

EURO-5K: When Does Domain Pretraining Matter? Benchmarking Transformers for EU Reporting Obligation Extraction

Model ReleasesDGX agent

arXiv:2606.02971v1 Announce Type: new Abstract: Extracting reporting obligations from EU legislation is critical for assessing and reducing regulatory reporting burden. However, distinguishing reporti

Evaluating and Calibrating LLM Confidence on Questions with Multiple Correct Answers

Model ReleasesDGX agent

arXiv:2602.07842v2 Announce Type: replace Abstract: Confidence calibration is essential for making large language models (LLMs) reliable, yet existing training-free methods have been primarily studied

← Previous
1…4546474849…129
Next →