AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlog
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,759 results
Model Releases

Beyond Literal Translation: Evaluating Cultural Effectiveness in Social Media UGC

DGX agent

arXiv:2605.25626v1 Announce Type: new Abstract: Social media platforms enable large-scale cross-lingual communication, but translating user-generated content (UGC) remains challenging due to its infor

model-releasesarxiv-cs-cl
26 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

Beyond the Target: From Imitation to Collaboration in Speculative Decoding

DGX agent

arXiv:2605.24793v1 Announce Type: new Abstract: Speculative decoding (SPD) accelerates large language model (LLM) inference by letting a smaller draft model propose multiple future tokens that are ver

safetyarxiv-cs-cl
26 May 2026
Safety

Can Large Language Models Resolve Semantic Discrepancy in Self-Destructive Subcultures? Evidence from Jirai Kei

DGX agent

arXiv:2601.05004v2 Announce Type: replace Abstract: Self-destructive behaviors are linked to complex psychological states and can be challenging to diagnose. These behaviors may be even harder to iden

safetyarxiv-cs-cl
26 May 2026
Model Releases

Can LoRA Fusion Support Cross-Domain Tasks in Cloud-Edge Collaboration?

DGX agent

arXiv:2605.23913v1 Announce Type: cross Abstract: Cloud-hosted large language models (LLMs) commonly rely on LoRA for domain adaptation, yet domain data are distributed across multiple edge devices an

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Clarification Is Not Enough: Post-Clarification Answering Remains the Bottleneck in Multi-Turn QA

DGX agent

arXiv:2605.25204v1 Announce Type: new Abstract: Pluralistic alignment requires systems to adapt to diverse user values, communication styles, and contextual assumptions. We believe that a foundational

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

CMAP: Cross-Modal Adaptive Prompting for Multi-Domain Task-Incremental Learning

DGX agent

arXiv:2605.25708v1 Announce Type: cross Abstract: Multi-domain task-incremental learning requires a model to sequentially acquire knowledge across visually diverse domains without forgetting prior tas

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions

DGX agent

arXiv:2605.24279v1 Announce Type: new Abstract: A frontier language model's acknowledged 'helpful programming assistant' persona does not survive long agentic-coding sessions in the deployment regime

model-releasesarxiv-cs-cl
26 May 2026
Agents

CP-Agent: A Calibrated Risk-Controlled Agent for Feedback-Driven Competitive Programming

DGX agent

arXiv:2605.24693v1 Announce Type: new Abstract: Large language models still struggle with contest-level programming, while many agentic remedies rely on massive inference-time sampling or expensive mu

agentsarxiv-cs-cl
26 May 2026
Safety

CRPO: Character-centric Group Relative Policy Optimization for Role-aware Reasoning in Role-playing Agents

DGX agent

arXiv:2605.25511v1 Announce Type: new Abstract: Recent advancements in Reinforcement Learning (RL), particularly Group Relative Policy Optimization (GRPO), have significantly enhanced the reasoning ca

safetyarxiv-cs-cl
26 May 2026
Model Releases

CSP-Atlas: Concept-Specific Neural Circuits in a Sparse Python Transformer

DGX agent

arXiv:2605.24603v1 Announce Type: new Abstract: A sparse 8-layer code transformer develops dedicated neural circuitry for every Python construct tested, and that circuitry is organised by a clean comp

model-releasesarxiv-cs-cl
26 May 2026
Research

CUNY at CLPsych 2026: A Pipeline Approach to Classification and Summarization of Mental Health Changes

DGX agent

arXiv:2605.24164v1 Announce Type: new Abstract: We describe our submission to the CLPsych~2026 Shared Task on capturing and characterizing mental health changes through social media timeline dynamics.

researcharxiv-cs-cl
26 May 2026
Model Releases

Decompose-and-Refine: Structured Legal Question Answering with Parametric Retrieval

DGX agent

arXiv:2605.24454v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong performance in the legal domain, demonstrating notable potential in Legal Question Answering (LQA). Howev

model-releasesarxiv-cs-cl
26 May 2026
Applications

DeIDClinic: A Risk-Aware Pseudonymization Framework for Clinical Text De-identification and Re-identification Risk Assessment

DGX agent

arXiv:2410.01648v2 Announce Type: replace Abstract: The increasing availability of sensitive textual data has created an urgent need for robust de-identification methods that enable compliant data sha

applicationsarxiv-cs-cl
26 May 2026
Safety

Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs

DGX agent

arXiv:2605.23975v1 Announce Type: new Abstract: Audio large language models (Audio LLMs) exhibit systematic failures in transcribing code-switching speech despite strong multilingual capabilities. Foc

safetyarxiv-cs-cl
26 May 2026
Model Releases

Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models

DGX agent

arXiv:2605.25189v1 Announce Type: cross Abstract: Reward hacking arises when a model improves a proxy reward by exploiting shortcuts rather than solving the intended task. We study this failure mode t

model-releasesarxiv-cs-cl
26 May 2026
Safety

Discovering Lexical Gaps Using Embeddings from Multilingual LLMs

DGX agent

arXiv:2605.24310v1 Announce Type: new Abstract: Lexical gaps are words that do not exist in certain languages. They pose challenges for building multilingual lexical resources, for machine translation

safetyarxiv-cs-cl
26 May 2026
Research

Distinguishing Right from Wrong in Debates: Attribution Analysis of Chinese Harmful Memes

DGX agent

arXiv:2605.24344v1 Announce Type: new Abstract: Research on harmful meme detection has garnered significant attention, resulting in the development of numerous datasets and methods. However, progress

researcharxiv-cs-cl
26 May 2026
Research

Does Continued Pretraining on a Learner Corpus Improve Automated Essay Scoring on English Proficiency Tests? Evidence from EFCAMDAT

DGX agent

arXiv:2605.25924v1 Announce Type: new Abstract: Recent automated essay scoring (AES) studies increasingly use pretrained transformer models, but these models are usually pretrained on general-domain E

researcharxiv-cs-cl
26 May 2026
Model Releases

Double Triangle Annotation: A Scalable Human-in-the-Loop Framework for High-Precision Historical Document Annotation

DGX agent

arXiv:2605.25781v1 Announce Type: new Abstract: Evaluating structured-information extraction from historical documents at scale requires high-precision ground-truth annotations, yet traditional manual

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

DRInQ: Evaluating Conversational Implicature with Controlled Context Variation

DGX agent

arXiv:2605.24267v1 Announce Type: new Abstract: Human conversation relies heavily on conversational implicature, in which speakers convey meanings that are suggested rather than explicitly stated. Alt

model-releasesarxiv-cs-cl
26 May 2026
Local Ai

DTO: a Differentiable Training Objective for Effective Counterfactual Story Rewriting

DGX agent

arXiv:2605.24885v1 Announce Type: new Abstract: Counterfactual story rewriting is a natural language processing task that requires updating an existing story to reflect a chosen alternative event, yet

local-aiarxiv-cs-cl
26 May 2026
Research

DUEL: Adversarial Self-Play for Multimodal Reasoning

DGX agent

arXiv:2605.24794v1 Announce Type: cross Abstract: Reinforcement learning (RL) has emerged as an effective paradigm for improving the reasoning capability of vision-language models (VLMs). However, RL-

researcharxiv-cs-cl
26 May 2026
Safety

DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning

DGX agent

arXiv:2605.25604v1 Announce Type: new Abstract: Reinforcement Learning has become a standard paradigm for aligning Large Language Models with human intent and task requirements. While Group Relative P

safetyarxiv-cs-cl
26 May 2026
Safety

ECHO: Terminal Agents Learn World Models for Free

DGX agent

arXiv:2605.24517v1 Announce Type: cross Abstract: CLI agents are the closest thing language models have to an embodied setting: the model emits commands, the terminal executes them, and the returned s

safetyarxiv-cs-cl
26 May 2026
Agents

EfficientGraph-RAG: Structured Retrieval-State Management for Cross-Task Retrieval-Augmented Generation

DGX agent

arXiv:2605.25379v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) has become the standard way to ground large language models in external knowledge, but many systems still organize

agentsarxiv-cs-cl
26 May 2026
Research

End-to-End Intracortical Speech Decoding from Neural Activity

DGX agent

arXiv:2605.24313v1 Announce Type: new Abstract: Current high-performing intracortical speech neuroprostheses achieve low word error rates but typically rely on external language models during inferenc

researcharxiv-cs-cl
26 May 2026
Research

Exploring Profiles of Cognitive Distortions Associated with Mental Health Disorders

DGX agent

arXiv:2605.24996v1 Announce Type: new Abstract: Cognitive distortions, distorted patterns of thinking, have been increasingly studied in computational mental health research. Although they are related

researcharxiv-cs-cl
26 May 2026
Safety

Faithful or Fabricated? A Causal Framework for Rationalization Bias in LLM Judges

DGX agent

arXiv:2605.23970v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as automatic judges for summarization and dialogue evaluation. Prior work has documented biases such

safetyarxiv-cs-cl
26 May 2026
Safety

Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning

DGX agent

arXiv:2605.24286v1 Announce Type: cross Abstract: Chain-of-thought (CoT) reasoning is useful for monitoring language models only when the reasoning trace faithfully reflects the computation that produ

safetyarxiv-cs-cl
26 May 2026
Model Releases

Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth

DGX agent

arXiv:2605.25052v1 Announce Type: new Abstract: Chains of thought (CoTs) have become central in interpreting and auditing behaviors of large language models. Yet growing evidence suggests that these t

model-releasesarxiv-cs-cl
26 May 2026
Research

Feature Resemblance: Towards a Theoretical Understanding of Analogical Reasoning in Transformers

DGX agent

arXiv:2603.05143v3 Announce Type: replace Abstract: Understanding reasoning in large language models is complicated by evaluations that conflate multiple reasoning types. We isolate analogical reasoni

researcharxiv-cs-cl
26 May 2026
Research

Forgotten Words: Benchmarking NeoBERT for Dementia Detection in Low-Resource Conversational Filipino and English Speech

DGX agent

arXiv:2605.26007v1 Announce Type: new Abstract: Dementia detection from spontaneous speech offers a scalable approach to cognitive screening, yet NLP systems remain predominantly English-centric. This

researcharxiv-cs-cl
26 May 2026
Model Releases

Found in Conversation: LLMs Teach Themselves to Close the Multi-Turn Gap

DGX agent

arXiv:2605.24432v1 Announce Type: new Abstract: Large Language Model (LLM) interactions are typically underspecified, with users clarifying all necessary details across multiple conversational turns.

model-releasesarxiv-cs-cl
26 May 2026
Safety

From Automation to Collaboration: Human-in-the-Loop Methods for Safe and Trustworthy NLP

DGX agent

arXiv:2605.25226v1 Announce Type: new Abstract: Large language models are widely deployed in high-stakes NLP tasks, yet risks such as bias, hallucination, adversarial vulnerability and unreliable gene

safetyarxiv-cs-cl
26 May 2026
Model Releases

From Facts to Insights: A Persona-Driven Dual Memory Framework and Dataset for Role-Playing Agents

DGX agent

arXiv:2605.25693v1 Announce Type: new Abstract: While role-playing agents excel in short-term interactions, long-term conversations overwhelm context windows, motivating external memory frameworks. Cu

model-releasesarxiv-cs-cl
26 May 2026
Safety

From Knowledge to Inference: Formalizing Specialized Public Health Reasoning on GlobalHealthAtlas

DGX agent

arXiv:2602.00491v2 Announce Type: replace Abstract: Public health reasoning requires population level inference grounded in scientific evidence, expert consensus, and safety constraints. However, it r

safetyarxiv-cs-cl
26 May 2026
Applications

Generating Legal Commentaries from Case Databases via Retrieval, Clustering, and Generation

DGX agent

arXiv:2605.24534v1 Announce Type: new Abstract: We present a fully automated pipeline that transforms large collections of court decisions into legal commentaries for statutes - without providing any

applicationsarxiv-cs-cl
26 May 2026
Research

GeoMathCode: Understanding Interleaved Math-Code Reasoning for Geometry Problem Solving

DGX agent

arXiv:2605.25384v1 Announce Type: new Abstract: Mathematical reasoning is a hallmark of human intelligence, requiring logical deduction, symbolic manipulation, and abstract thinking. Recent multimodal

researcharxiv-cs-cl
26 May 2026
Safety

GeoSVG-RL: Geometry-Aware Reinforcement Learning for Layout-Constrained Text-to-SVG Diagram Generation

DGX agent

arXiv:2605.25447v1 Announce Type: new Abstract: Generating structured, editable diagrams remains a significant challenge for contemporary large language models, despite their proficiency in general-pu

safetyarxiv-cs-cl
26 May 2026
Model Releases

GroupTravelBench: Benchmarking LLM Agents on Multi-Person Travel Planning

DGX agent

arXiv:2605.25200v1 Announce Type: new Abstract: Travel planning is a realistic task for evaluating the planning and tool-use abilities of LLM agents. However, existing benchmarks typically assume only

model-releasesarxiv-cs-cl
26 May 2026
Hardware

H^{2}MT: Semantic Hierarchy-Aware Hierarchical Memory Transformer

DGX agent

arXiv:2605.24930v1 Announce Type: new Abstract: Transformer-based LLMs achieve strong results on many language tasks; however, long inputs remain challenging because context windows are finite, and pr

hardwarearxiv-cs-cl
26 May 2026
Safety

Harmony in Diversity: Multi-domain Contrastive Policy Optimization for Large Reasoning Models

DGX agent

arXiv:2605.25443v1 Announce Type: new Abstract: Post-training has significantly enhanced the reasoning capability of Large Reasoning Models (LRMs), especially with Reinforcement Learning (RL) like Gro

safetyarxiv-cs-cl
26 May 2026
Local Ai

Hierarchical Local-Global Transformer for Temporal Sentence Grounding

DGX agent

arXiv:2208.14882v2 Announce Type: replace-cross Abstract: This paper studies the multimedia problem of temporal sentence grounding (TSG), which aims to accurately determine the specific video segment

local-aiarxiv-cs-cl
26 May 2026
Model Releases

HiMed: Incentivizing Hindi Reasoning in Medical LLMs

DGX agent

arXiv:2605.24635v1 Announce Type: new Abstract: Medical large language models hold promise for reducing healthcare disparities, yet Hindi remains severely underrepresented. While medical LLMs excel in

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework

DGX agent

arXiv:2507.19219v2 Announce Type: replace Abstract: Overestimation in evaluating large language models (LLMs) has become an increasing concern. Due to the contamination of public benchmarks or imbalan

model-releasesarxiv-cs-cl
26 May 2026
Safety

How Much Structure Do LLMs Need? Evaluating LLMs for Bibliometric Cluster Description

DGX agent

arXiv:2605.24351v1 Announce Type: new Abstract: Large language models (LLMs) can support scientific literature synthesis, but remain prone to hallucinated references, uneven coverage, and weakly groun

safetyarxiv-cs-cl
26 May 2026
Agents

HyLaT: Efficient Multi-Agent Communication via Hybrid Latent-Text Protocol

DGX agent

arXiv:2605.25421v1 Announce Type: new Abstract: Communication protocol design is a central challenge in large language model-based multi-agent systems. Existing single-channel approaches face an inher

agentsarxiv-cs-cl
26 May 2026
Safety

Improving the Completeness and Comparability of Segment Disclosures: A Large Language Model Approach

DGX agent

arXiv:2605.23924v1 Announce Type: new Abstract: Segment-level disclosures are a central component of financial reporting, providing insight into firms' internal organization and the allocation of econ

safetyarxiv-cs-cl
26 May 2026
← Previous
1…7677787980…162
Next →