AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
Model Releases

Video2LoRA: Parametric Video Internalization for Vision-Language Models

DGX agent

arXiv:2606.04351v1 Announce Type: cross Abstract: Processing video in vision-language models is expensive: each frame occupies hundreds of tokens, and inference cost scales with every frame and every

model-releasesarxiv-cs-cl
4 Jun 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

WETBench: A Benchmark for Detecting Task-Specific Machine-Generated Text on Wikipedia

DGX agent

arXiv:2507.03373v2 Announce Type: replace Abstract: Given Wikipedia's role as a trusted source of high-quality, reliable content, concerns are growing about the proliferation of low-quality machine-ge

model-releasesarxiv-cs-cl
4 Jun 2026
Research

When Clients Stop Following: A Cognitive Conceptualization Diagram-driven Framework for Strategic Counseling

DGX agent

arXiv:2606.04389v1 Announce Type: new Abstract: Large Language Models (LLMs) show promise in psychological counseling, yet existing benchmarks rely heavily on highly cooperative simulated clients. We

researcharxiv-cs-cl
4 Jun 2026
Research

When Retrieval Doesn't Help: A Large-Scale Study of Biomedical RAG

DGX agent

arXiv:2606.04127v1 Announce Type: new Abstract: Medical question answering is a high-stakes setting where factual errors can have serious consequences. Retrieval-augmented generation (RAG) is widely v

researcharxiv-cs-cl
4 Jun 2026
Research

A cross-domain tropical species dataset with Chinese vernacular names and CITES source links

DGX agent

arXiv:2606.03156v1 Announce Type: new Abstract: We describe a versioned cross-domain dataset of 410,499 active tropical species (working snapshot 2026-04-20) spanning three applied subdomains -- tropi

researcharxiv-cs-cl
3 Jun 2026
Tutorials

A Locally Deployed RAG-Based Academic Advising System for Course Selection

DGX agent

arXiv:2606.02983v1 Announce Type: new Abstract: The correct sequence of courses in the curriculum based on prerequisites between courses is of great importance for students to develop their knowledge

tutorialsarxiv-cs-cl
3 Jun 2026
Safety

A Pocket Offline Model for Simultaneous Speech Translation as CUNI Submission to IWSLT 2026

DGX agent

arXiv:2606.03948v1 Announce Type: new Abstract: We implement simultaneous translation capability with the offline direct speech-to-text translation model Canary, using the state-of-the-art policy Alig

safetyarxiv-cs-cl
3 Jun 2026
Local Ai

ARBOR: Online Process Rewards via a Reusable Rubric Buffer for Search Agents

DGX agent

arXiv:2606.03239v1 Announce Type: new Abstract: LLM-based search agents are trained predominantly with outcome-only reward, leaving the search process itself unsupervised. This signal degenerates on o

local-aiarxiv-cs-cl
3 Jun 2026
Applications

Assessing Pause Thresholds for empirical Translation Process Research

DGX agent

arXiv:2604.01410v2 Announce Type: replace Abstract: Text production (and translations) proceeds in the form of stretches of typing, interrupted by keystroke pauses. It is often assumed that fast typin

applicationsarxiv-cs-cl
3 Jun 2026
Research

AutoTail-BSFGM: Class-Balance-Aware Fine-Tuning for Chinese Scholarly Text Classification

DGX agent

arXiv:2606.03576v1 Announce Type: new Abstract: Scholarly text classification supports literature organization, subject indexing, and research intelligence, but Chinese scholarly corpora often contain

researcharxiv-cs-cl
3 Jun 2026
Safety

Backdoor Unlearning Generalization: A Path Toward the Removal of Unknown Triggers in LLMs

DGX agent

arXiv:2606.03785v1 Announce Type: new Abstract: Backdoor attacks in Large Language Models (LLMs) are a growing security concern, where models can generate adversary-chosen content. Existing defenses t

safetyarxiv-cs-cl
3 Jun 2026
Applications

Benchmarking Speech-to-Speech Translation Models

DGX agent

arXiv:2606.03241v1 Announce Type: new Abstract: Speech-to-speech translation (S2ST) has advanced rapidly, but offline evaluation lacks a unified protocol: studies report non-overlapping metric subsets

applicationsarxiv-cs-cl
3 Jun 2026
Model Releases

Beyond Ideal Instruction: A Comprehensive Framework for Evaluating LLMs in Realistic Interactions

DGX agent

arXiv:2606.03318v1 Announce Type: new Abstract: Despite great advances in tool-use capabilities of large language models (LLMs), existing evaluation benchmarks struggle to fully align with real-world

model-releasesarxiv-cs-cl
3 Jun 2026
Research

Beyond Semantics: Modeling Factual and Affective Perceptual Experiences from Vision-Language Data

DGX agent

arXiv:2606.03345v1 Announce Type: cross Abstract: We present P-Topics (Perception Topics) modeling, a novel problem for understanding how images are perceived affectively and across cultures. The goal

researcharxiv-cs-cl
3 Jun 2026
Research

Beyond the Literal: Decomposing Pragmatic Intent in Multimodal Meme Understanding

DGX agent

arXiv:2606.03604v1 Announce Type: new Abstract: When asked what a meme or sarcastic post means, Large Vision Language Models (LVLMs) tend to describe what the image shows rather than what the author i

researcharxiv-cs-cl
3 Jun 2026
Research

Beyond 'To whom it may concern': Tailoring Machine Translation to Audience and Intent

DGX agent

arXiv:2606.03259v1 Announce Type: new Abstract: Translation quality depends on purpose: the same source text demands different translations depending on audience, tone, and communicative intent. Yet M

researcharxiv-cs-cl
3 Jun 2026
Safety

Breaking the Self-Confirming Loop: Diagnosing and Mitigating Systemic Reward Bias in Self-Rewarding RL

DGX agent

arXiv:2510.08977v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) efficiently scales the reasoning ability of large language models (LLMs) but is bottlene

safetyarxiv-cs-cl
3 Jun 2026
Model Releases

Can Factual Opinions Be Edited (Manipulated) in Large Language Models?

DGX agent

arXiv:2606.03096v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly integrated into various domains, making knowledge editing techniques crucial yet potentially hazardous. Cu

model-releasesarxiv-cs-cl
3 Jun 2026
Research

Can LLM Rerankers Predict Their Own Ranking Performance?

DGX agent

arXiv:2606.03535v1 Announce Type: cross Abstract: Retrieval effectiveness varies substantially across queries, making it important to estimate ranking quality before relevance judgments are available.

researcharxiv-cs-cl
3 Jun 2026
Model Releases

Can Structural Cues Save LLMs? Evaluating Language Models in Massive Document Streams

DGX agent

arXiv:2603.19250v2 Announce Type: replace Abstract: Evaluating language models in streaming environments is critical, yet underexplored. Existing benchmarks either focus on single complex events or pr

model-releasesarxiv-cs-cl
3 Jun 2026
Model Releases

CAPER: Clause-Aligned Process Supervision for Text-to-SQL

DGX agent

arXiv:2606.03327v1 Announce Type: cross Abstract: Text-to-SQL systems are typically evaluated by query-level execution correctness, but this terminal signal provides little guidance about which interm

model-releasesarxiv-cs-cl
3 Jun 2026
Model Releases

Chatbots Output Meaningful (but Problematic) Language

DGX agent

arXiv:2606.02973v1 Announce Type: new Abstract: Are utterances by AI chatbots meaningful? Concretely, if a user asks, say, Anthropic's agent Claude, 'What is the capital of Spain?' and Claude answers,

model-releasesarxiv-cs-cl
3 Jun 2026
Safety

Coherence Maximization Improves Pluralistic Alignment

DGX agent

arXiv:2606.03110v1 Announce Type: new Abstract: Aligning AI systems with diverse human values requires value specifications grounded in concrete examples, but generating such examples without extensiv

safetyarxiv-cs-cl
3 Jun 2026
Applications

Core-based Hierarchies for Efficient GraphRAG

DGX agent

arXiv:2603.05207v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) enhances large language models by incorporating external knowledge. However, existing vector-based method

applicationsarxiv-cs-cl
3 Jun 2026
Model Releases

CourseTimeQA: A Lecture-Video Benchmark and a Latency-Constrained Cross-Modal Fusion Method for Timestamped QA

DGX agent

arXiv:2512.00360v2 Announce Type: replace Abstract: We study timestamped question answering over educational lecture videos under a single-GPU latency/memory budget. Given a natural-language query, th

model-releasesarxiv-cs-cl
3 Jun 2026
Research

CTR-Sink: Attention Sink for Language Models in Click-Through Rate Prediction

DGX agent

arXiv:2508.03668v2 Announce Type: replace Abstract: Click-Through Rate (CTR) prediction, a core task in recommendation systems, estimates user click likelihood using historical behavioral data. Modeli

researcharxiv-cs-cl
3 Jun 2026
Local Ai

DMT-CBT: Longitudinal Therapeutic State Modeling for CBT Counseling

DGX agent

arXiv:2606.03132v1 Announce Type: new Abstract: Large language models (LLMs) have shown growing potential for Cognitive Behavioral Therapy (CBT) counseling. However, most existing approaches still for

local-aiarxiv-cs-cl
3 Jun 2026
Model Releases

Do Value Vectors in Deep Layers Need Context from the Residual Stream?

DGX agent

arXiv:2606.02780v1 Announce Type: new Abstract: The success of the transformer architecture as the backbone of modern LLMs is in large part due to its use of attention layers. An attention layer follo

model-releasesarxiv-cs-cl
3 Jun 2026
Applications

Does Language Shift Break Medical Vision-Language Models? Indonesian Radiology Visual Question Answering Case Study

DGX agent

arXiv:2606.03693v1 Announce Type: new Abstract: Medical Vision-Language Models (VLMs) are typically evaluated on English radiology visual question answering benchmarks, leaving their robustness under

applicationsarxiv-cs-cl
3 Jun 2026
Model Releases

Don't Forget Your Embeddings: Robust Knowledge Erasure via Precise Editing of Embeddings

DGX agent

arXiv:2606.03695v1 Announce Type: new Abstract: As language models are increasingly deployed in real-world applications, the ability to erase specific knowledge from them becomes critical for safety a

model-releasesarxiv-cs-cl
3 Jun 2026
Safety

Dynamic Short Convolutions Improve Transformers

DGX agent

arXiv:2606.03825v1 Announce Type: cross Abstract: Transformers have become the dominant architecture for large language models, largely due to the scalability and flexibility of attention, feed-forwar

safetyarxiv-cs-cl
3 Jun 2026
Research

Enhancing Paraphrase Type Generation: The Impact of DPO and RLHF Evaluated with Human-Ranked Data

DGX agent

arXiv:2506.02018v2 Announce Type: replace Abstract: Paraphrasing re-expresses meaning to enhance applications like text simplification, machine translation, and question-answering. Specific paraphrase

researcharxiv-cs-cl
3 Jun 2026
Agents

Entropy Gate: Entropy Quenching for Near-Lossless Token Compression in LLM Pipelines

DGX agent

arXiv:2606.03739v1 Announce Type: new Abstract: LLM pipelines waste substantial token budgets on low-information content: repeated context, verbose responses, and redundant boilerplate. We introduce E

agentsarxiv-cs-cl
3 Jun 2026
Model Releases

EntSQL: A Benchmark for Grounding Text-to-SQL in Long-Context Enterprise Knowledge

DGX agent

arXiv:2606.03363v1 Announce Type: new Abstract: Text-to-SQL enables natural language access to databases, and recent LLMs have substantially advanced its capabilities. Existing benchmarks such as Spid

model-releasesarxiv-cs-cl
3 Jun 2026
Model Releases

EURO-5K: When Does Domain Pretraining Matter? Benchmarking Transformers for EU Reporting Obligation Extraction

DGX agent

arXiv:2606.02971v1 Announce Type: new Abstract: Extracting reporting obligations from EU legislation is critical for assessing and reducing regulatory reporting burden. However, distinguishing reporti

model-releasesarxiv-cs-cl
3 Jun 2026
Model Releases

Evaluating and Calibrating LLM Confidence on Questions with Multiple Correct Answers

DGX agent

arXiv:2602.07842v2 Announce Type: replace Abstract: Confidence calibration is essential for making large language models (LLMs) reliable, yet existing training-free methods have been primarily studied

model-releasesarxiv-cs-cl
3 Jun 2026
Model Releases

Experience-Driven Dynamic Exits for LLMs with Reinforcement Learning

DGX agent

arXiv:2606.03113v1 Announce Type: new Abstract: Large Language Models suffer from slow autoregressive inference. While self-speculative decoding accelerates this process, its efficiency is hampered by

model-releasesarxiv-cs-cl
3 Jun 2026
Local Ai

Expert-Aware Causal Tracing of Factual Recall in Sparse MoE Language Models

DGX agent

arXiv:2606.03780v1 Announce Type: new Abstract: Causal tracing of factual recall has been studied predominantly in dense transformer language models, where interventions localize information flow to l

local-aiarxiv-cs-cl
3 Jun 2026
Safety

Exploring Adversarial Robustness and Safety Alignment in Multilingual Multi-Modal Large Language Models

DGX agent

arXiv:2606.03793v1 Announce Type: new Abstract: Multimodal Large Language Models integrate visual perception into language reasoning, introducing a continuous attack surface susceptible to adversarial

safetyarxiv-cs-cl
3 Jun 2026
Local Ai

FederatedSkill: Federated Learning for Agentic Skill Evolution

DGX agent

arXiv:2606.03143v1 Announce Type: cross Abstract: Modern LLM agents increasingly rely on skill libraries to handle complex tasks, making skill evolution a primary driver of self-improvement. However,

local-aiarxiv-cs-cl
3 Jun 2026
Local Ai

Framing Migration News with LLMs: Structured CoT as a Support for Human Interpretation

DGX agent

arXiv:2606.03761v1 Announce Type: new Abstract: Frame analysis of migration news is a socially consequential task: media scholars and researchers who study how migration is narrated need tools that ar

local-aiarxiv-cs-cl
3 Jun 2026
Model Releases

From Script to Semantics: Prompting Strategies for African NLI

DGX agent

arXiv:2606.03304v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly evaluated in multilingual settings, yet their inference behavior in low-resource African languages remains

model-releasesarxiv-cs-cl
3 Jun 2026
Model Releases

G^2C-MT: Graph-Guided Context Selection for Document-Level Machine Translation

DGX agent

arXiv:2606.03078v1 Announce Type: new Abstract: Effective document-level machine translation (DocMT) requires capturing long-range discourse dependencies. Recent work has explored retrieval-based and

model-releasesarxiv-cs-cl
3 Jun 2026
Safety

GLINT: Sparsely Gated Vision-Language Alignment for Fine-Grained Radiology Representations

DGX agent

arXiv:2606.03180v1 Announce Type: cross Abstract: Vision-language models (VLMs) for radiology have emerged as a scalable paradigm by leveraging image-report pairs naturally produced in clinical workfl

safetyarxiv-cs-cl
3 Jun 2026
Model Releases

Greener Than Humans? Environmental Attitudes in Large Language Models

DGX agent

arXiv:2606.02741v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in sustainability-related decision support, reporting, and public communication, yet little systemati

model-releasesarxiv-cs-cl
3 Jun 2026
Model Releases

Hallucination Is Linearly Decodable from Mid-Layer Hidden States in Quantized LLMs

DGX agent

arXiv:2606.02628v1 Announce Type: cross Abstract: We investigate whether open-source LLMs encode a linearly separable truthfulness signal in their hidden states, and at which network depth this signal

model-releasesarxiv-cs-cl
3 Jun 2026
Safety

Hint-Guided Diversified Policy Optimization for LLM Reasoning

DGX agent

arXiv:2606.03021v1 Announce Type: new Abstract: Recent developments in Large Language Models (LLMs) have showcased impressive reasoning capabilities, with Reinforcement Learning with Verifiable Reward

safetyarxiv-cs-cl
3 Jun 2026
Research

HybridThinker: Efficient Chain-of-Thought Reasoning via Compressed Memory and Transient Thought Steps

DGX agent

arXiv:2606.03768v1 Announce Type: new Abstract: Extended chain-of-thought (CoT) traces improve LLM reasoning but incur substantial computational and memory costs. While existing CoT compression method

researcharxiv-cs-cl
3 Jun 2026
← Previous
1…5758596061…161
Next →