AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
Safety

Alignment Risks from Capability-Seeking RL Training

DGX agent

arXiv:2602.12124v2 Announce Type: replace-cross Abstract: While most AI alignment research focuses on preventing models from generating explicitly harmful content, a more subtle risk arises from capab

safetyarxiv-cs-cl
5 Jun 2026
Model Releases
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Almieyar-Oryx-BloomBench: A Bilingual Multimodal Benchmark for Cognitively Informed Evaluation of Vision-Language Models

DGX agent

arXiv:2606.05531v1 Announce Type: cross Abstract: Despite the rapid progress of Vision-Language Models (VLMs), the field lacks benchmarks that rigorously diagnose their true reasoning abilities and ch

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

An Embarrassingly Simple Detector for Model Extraction Attacks in Large Language Model API Traffic

DGX agent

arXiv:2606.05725v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed through hosted APIs, making model extraction a practical threat to model ownership and service

model-releasesarxiv-cs-cl
5 Jun 2026
Research

An ERP Study on Recursive Locative Processing in Mandarin-Speaking Children with Autism

DGX agent

arXiv:2606.05620v1 Announce Type: new Abstract: Recursion enables the generation of hierarchical linguistic structures but imposes substantial processing demands during real-time comprehension. While

researcharxiv-cs-cl
5 Jun 2026
Safety

Analysis of the Neglect-Zero Effect in Large Language Models

DGX agent

arXiv:2606.05864v1 Announce Type: new Abstract: We investigate the extent to which the language processing of LLMs resembles human cognitive processes, focusing on a human cognitive bias called the ex

safetyarxiv-cs-cl
5 Jun 2026
Model Releases

ArcANE: Do Role-Playing Language Agents Stay in Character at the Right Time?

DGX agent

arXiv:2606.05553v1 Announce Type: new Abstract: Role-playing language agents (RPLAs) should play characters whose values and behavior evolve as the story progresses, not maintain a fixed persona. Exis

model-releasesarxiv-cs-cl
5 Jun 2026
Agents

Ask or Assume? Uncertainty-Aware Clarification-Seeking in Coding Agents

DGX agent

arXiv:2603.26233v2 Announce Type: replace Abstract: As Large Language Model (LLM) agents are increasingly deployed in open-ended domains like software engineering, they frequently encounter underspeci

agentsarxiv-cs-cl
5 Jun 2026
Model Releases

Asuka-Bench: Benchmarking Code Agents on Underspecified User Intent and Multi-Round Refinement

DGX agent

arXiv:2606.05920v1 Announce Type: cross Abstract: Existing code-generation benchmarks score a single mapping from a complete prompt to a one-shot output. However, real web development is different. Us

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

AURA: Intent-Directed Probing for Implicit-Need Surfacing in Situated LLM Agents

DGX agent

arXiv:2606.05557v1 Announce Type: new Abstract: A situated query like 'where is Lin Wei?' often encodes more than its literal content: the user may also want to know whether Lin Wei is free, in a good

model-releasesarxiv-cs-cl
5 Jun 2026
Research

Automatic Labelling of Speech Translation Errors

DGX agent

arXiv:2606.06047v1 Announce Type: new Abstract: Errors in speech translations reduce trustworthiness of Speech Translation (ST) systems and can have serious consequences. Yet currently there is no est

researcharxiv-cs-cl
5 Jun 2026
Model Releases

Benchmarking Open-Source Layout Detection Models for Data Snapshot Extraction from Institutional Documents

DGX agent

arXiv:2606.06242v1 Announce Type: new Abstract: Institutional documents contain substantial amounts of operational and analytical information embedded within figures and tables. Current approaches for

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

Better Literary Translation: A Multi-Aspect Data Generation and LLM Training Approach

DGX agent

arXiv:2606.05924v1 Announce Type: new Abstract: Literary translation poses unique challenges due to the scarcity of high-quality annotated data and the need to balance expression fluency with literary

model-releasesarxiv-cs-cl
5 Jun 2026
Safety

Beyond Alignment: Value Diversity as a Collective Property in Multicultural Agent Systems

DGX agent

arXiv:2606.05985v1 Announce Type: new Abstract: Multicultural multi-agent systems are increasingly deployed in globally diverse settings, where different agents are grounded in different cultural back

safetyarxiv-cs-cl
5 Jun 2026
Safety

Beyond tokens: a unified framework for latent communication in LLM-based multi-agent systems

DGX agent

arXiv:2606.05711v1 Announce Type: new Abstract: Multi-agent systems built on large language models (LLMs) have become a prevailing paradigm for tackling complex reasoning, planning, and tool-use tasks

safetyarxiv-cs-cl
5 Jun 2026
Applications

Bootstrapping Semantic Layer from Execution for Text-to-SQL

DGX agent

arXiv:2606.05634v1 Announce Type: new Abstract: Real-world text-to-SQL is often under-specified until user phrases are grounded in how the database stores values. Prior work attempts to address this b

applicationsarxiv-cs-cl
5 Jun 2026
Research

CaliDist: Calibrating Large Language Models via Behavioral Robustness to Distraction

DGX agent

arXiv:2606.05799v1 Announce Type: cross Abstract: Existing calibration methods for Large Language Models (LLMs) often overlook a critical dimension of trustworthiness: a model's {em behavioral robustn

researcharxiv-cs-cl
5 Jun 2026
Model Releases

Can LLMs Be Constrained to the Past? Improving Knowledge Cutoff through Recall-Based Prompting

DGX agent

arXiv:2606.05804v1 Announce Type: new Abstract: Prompted knowledge cutoff instructs a large language model (LLM) to act as if information beyond a specified cutoff date were unavailable. However, prio

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

CHALIS: A Challenge Dataset for Language Identification in Difficult Scenarios

DGX agent

arXiv:2606.06088v1 Announce Type: new Abstract: We present CHALIS (Challenging Language Identification Samples), a new benchmark dataset explicitly designed to address difficult cases in language iden

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

Channel-Wise Mixed-Precision Quantization for Large Language Models

DGX agent

arXiv:2410.13056v4 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated remarkable success across a wide range of language tasks, but their deployment on edge devices remain

model-releasesarxiv-cs-cl
5 Jun 2026
Research

ChartAttack: Testing the Vulnerability of LLMs to Malicious Prompting in Chart Generation

DGX agent

arXiv:2601.12983v3 Announce Type: replace Abstract: Multimodal large language models (MLLMs) are increasingly used to automate chart generation from data tables, improving analysis and reporting effic

researcharxiv-cs-cl
5 Jun 2026
Safety

CHASE: Adversarial Red-Blue Teaming for Improving LLM Safety using Reinforcement Learning

DGX agent

arXiv:2606.05523v1 Announce Type: new Abstract: Despite advances in safety alignment, prompt-rewriting attacks such as persona modulation, fictional framing and persuasion-based reformulation, can byp

safetyarxiv-cs-cl
5 Jun 2026
Safety

'Chi nas dal soch el sent de legn' -- Auditing Text Corpora for Lombard

DGX agent

arXiv:2606.06349v1 Announce Type: new Abstract: Several of the world's languages are still under-resourced in terms of Natural Language Processing (NLP) tools. This is mostly due to the lack of high-q

safetyarxiv-cs-cl
5 Jun 2026
Model Releases

CLASH: Evaluating Language Models on Judging High-Stakes Dilemmas from Multiple Perspectives

DGX agent

arXiv:2504.10823v4 Announce Type: replace Abstract: Navigating dilemmas involving conflicting values is challenging even for humans in high-stakes domains, let alone for AI, yet prior work has been li

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

CLFEC: A New Task for Unified Linguistic and Factual Error Correction in paragraph-level Chinese Professional Writing

DGX agent

arXiv:2602.23845v2 Announce Type: replace Abstract: Chinese text correction has traditionally focused on spelling and grammar, while factual error correction is usually treated separately. However, in

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

Code2LoRA: Hypernetwork-Generated Adapters for Code Language Models under Software Evolution

DGX agent

arXiv:2606.06492v1 Announce Type: cross Abstract: Code language models need repository-level context to resolve imports, APIs, and project conventions. Existing methods inject this knowledge as long i

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

Coding with 'Enemy': Can Human Developers Detect AI Agent Sabotage?

DGX agent

arXiv:2606.05647v1 Announce Type: cross Abstract: AI coding agents are increasingly embedded in real-world software development, collaborating with human developers while gaining broader access to cod

model-releasesarxiv-cs-cl
5 Jun 2026
Research

ColBERTSaR: Sparsified ColBERT Index via Product Quantization

DGX agent

arXiv:2606.05568v1 Announce Type: cross Abstract: While ColBERT is an effective neural retrieval architecture, it requires a heavy index structure to support candidate set retrieval based on approxima

researcharxiv-cs-cl
5 Jun 2026
Model Releases

CollabBench: Benchmarking and Unleashing Collaborative Ability of LLMs with Diverse Players via Proactive Engagement

DGX agent

arXiv:2606.05793v1 Announce Type: new Abstract: While LLM-based agents excel at individual tasks, effective collaboration with realistic human partners remains challenging. Most of the existing conver

model-releasesarxiv-cs-cl
5 Jun 2026
Agents

CollabSim: A CSCW-Grounded Methodology for Investigating Collaborative Competence of LLM Agents through Controlled Multi-Agent Experiments

DGX agent

arXiv:2606.06399v1 Announce Type: new Abstract: Multi-agent systems (MAS) built on large language models have shown growing promise, with their effectiveness resting on agents' ability to coordinate t

agentsarxiv-cs-cl
5 Jun 2026
Model Releases

CoMoL: Efficient Mixture of LoRA Experts via Dynamic Core Space Merging

DGX agent

arXiv:2603.00573v2 Announce Type: replace Abstract: Large language models (LLMs) achieve remarkable performance on diverse downstream and domain-specific tasks via parameter-efficient fine-tuning (PEF

model-releasesarxiv-cs-cl
5 Jun 2026
Research

ComplexityMT: Benchmarking the Interaction Between Text Complexity and Machine Translation

DGX agent

arXiv:2606.05421v1 Announce Type: new Abstract: When a text is translated, does the translation retain the complexity of the original? We introduce ComplexityMT, a new challenge for assessing how text

researcharxiv-cs-cl
5 Jun 2026
Model Releases

Compress-Distill: Reasoning Trace Compression for Efficient Knowledge Distillation

DGX agent

arXiv:2606.05988v1 Announce Type: cross Abstract: Reasoning models produce long chain-of-thought traces that are costly to distill and encourage verbose student outputs. We study post-hoc compression

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

Contextualized Prompting For Stance Detection On Social Media

DGX agent

arXiv:2606.06022v1 Announce Type: new Abstract: Stance detection on social media is challenging due to short, noisy, and context-dependent language. While large language models (LLMs) show zero-shot g

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

Continual Learning Bench: Evaluating Frontier AI Systems in Real-World Stateful Environments

DGX agent

arXiv:2606.05661v1 Announce Type: cross Abstract: Continual learning, the ability of AI systems to improve through sequential experience, has attracted substantial interest, but no high-quality benchm

model-releasesarxiv-cs-cl
5 Jun 2026
Research

Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering

DGX agent

arXiv:2510.05709v2 Announce Type: replace-cross Abstract: LLM benchmarking metrics often misstate performance and uncertainty as they rely on two assumptions that frequently do not hold in practice: (

researcharxiv-cs-cl
5 Jun 2026
Research

CoT-Space: A Theoretical Framework for Internal Slow-Thinking via Reinforcement Learning

DGX agent

arXiv:2509.04027v3 Announce Type: replace-cross Abstract: Test-time scaling, primarily manifested through multi-step Chain-of-Thought (CoT) reasoning via Reinforcement Learning (RL), has emerged as a

researcharxiv-cs-cl
5 Jun 2026
Research

Decomposing Factual Sycophancy in Language Models: How Size and Instruction Tuning Shape Robustness

DGX agent

arXiv:2606.06306v1 Announce Type: new Abstract: Factual sycophancy occurs when a language model abandons a correct, verifiable answer under social pressure. Because a flip occurs only when pressure to

researcharxiv-cs-cl
5 Jun 2026
Model Releases

Dense Contexts Are Hard Contexts: Lexical Density Limits Effective Context in LLMs

DGX agent

arXiv:2606.06203v1 Announce Type: new Abstract: Input length and the position of relevant information are widely cited as the primary causes of degraded LLM long-context performance. Here, we study le

model-releasesarxiv-cs-cl
5 Jun 2026
Research

DiG-Plan: Mitigating Early Commitment for Tool-Graph Planning via Diffusion Guidance

DGX agent

arXiv:2606.05728v1 Announce Type: cross Abstract: Generating executable tool plans requires selecting appropriate subsets from tool libraries, a combinatorial search problem with an exponentially larg

researcharxiv-cs-cl
5 Jun 2026
Model Releases

Do MLLMs Capture How Interfaces Guide User Behavior? A Benchmark for Multimodal UI/UX Design Understanding

DGX agent

arXiv:2505.05026v5 Announce Type: replace Abstract: User interface (UI) design goes beyond visuals to shape user experience (UX), underscoring the shift toward UI/UX as a unified concept. While recent

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

DocHop-QA: Towards Multi-Hop Reasoning over Multimodal Document Collections

DGX agent

arXiv:2508.15851v2 Announce Type: replace Abstract: Despite rapid progress in large language models (LLMs), current QA benchmarks still overlook the core challenge of real-world scientific information

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

Domain-Aware Mispronunciation Detection and Diagnosis Using Language-Specific Statistical Graphs

DGX agent

arXiv:2606.05569v1 Announce Type: new Abstract: Mispronunciation Detection and Diagnosis (MDD) has gained increasing importance in computer-assisted language learning and speech technology in recent y

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

Domain-Conditioned Safety in Frontier Computer-Using Agents: A 793-Episode Browser Benchmark, a Coding-Domain Cross-Reference, and a Reproducibility Audit of Recent Red-Teaming

DGX agent

arXiv:2606.05233v1 Announce Type: cross Abstract: Recent computer-using-agent (CUA) red-teaming papers report prompt-injection attack success rates (ASR) of 42-98%, but these headline numbers cluster

model-releasesarxiv-cs-cl
5 Jun 2026
Research

Dynamic Thinking-Token Selection for Efficient Reasoning in Large Reasoning Models

DGX agent

arXiv:2601.18383v2 Announce Type: replace-cross Abstract: Large Reasoning Models (LRMs) excel at solving complex problems by explicitly generating a reasoning trace before deriving the final answer. H

researcharxiv-cs-cl
5 Jun 2026
Applications

EDIT: Evidence-Diagnosed Intervention Training for Rule-Faithful LLM Grading

DGX agent

arXiv:2606.06350v1 Announce Type: new Abstract: Reliable rubric grading requires more than accurate score prediction. Each judgement must be grounded in the mark scheme and evidence from the student a

applicationsarxiv-cs-cl
5 Jun 2026
Model Releases

Efficient Punctuation Restoration via Weighted Lookahead Scoring Method for Streaming ASR Systems

DGX agent

arXiv:2606.05179v1 Announce Type: new Abstract: Punctuation restoration improves ASR (Automatic Speech Recognition) readability. However streaming ASR requires online decisions with limited future con

model-releasesarxiv-cs-cl
5 Jun 2026
Agents

EGTR-Review: Efficient Evidence-Grounded Scientific Peer Review Generation via Multi-Agent Teacher Distillation

DGX agent

arXiv:2606.06025v1 Announce Type: new Abstract: Scientific peer review generation has attracted increasing attention for reducing reviewing burdens and providing timely feedback. However, existing Lar

agentsarxiv-cs-cl
5 Jun 2026
Safety

EMBER: Efficient Memory via Budgeted Evidence Retention for Long-Horizon Agents

DGX agent

arXiv:2606.05894v1 Announce Type: new Abstract: Long-horizon agents can archive large histories, but future answers still incur retrieval, rereading, and context costs. When retained memory misses ans

safetyarxiv-cs-cl
5 Jun 2026
← Previous
1…5152535455…161
Next →