AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
8 Jun 2026

PromptPrint: Behavioral Biometrics Through Natural Language Prompting in LLMs

Model ReleasesDGX agent

arXiv:2606.06755v1 Announce Type: new Abstract: Authorship attribution research has traditionally focused on long-form, expressive texts; however, interactions with large language models (LLMs) are ty

Quantifying Media Representation Dynamics Across 25 Years of News Reporting on Policing-related Deaths

ResearchDGX agent

arXiv:2606.06812v1 Announce Type: new Abstract: We perform the largest known computational analysis of Canadian news narratives about police-involved deaths, spanning 4,000 articles from the last quar

RASFT: Rollout-Adaptive Supervised Fine-Tuning for Reasoning

SafetyDGX agent

arXiv:2606.07006v1 Announce Type: cross Abstract: Supervised fine-tuning (SFT) is a prevailing method for adapting large language models to reasoning tasks by imitating offline expert demonstrations,


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

RECAP: Regression Evaluation for Continual Adaptation of Prompts

Model ReleasesDGX agent

arXiv:2606.06698v1 Announce Type: cross Abstract: Production agentic systems routinely face evolving constraints and must comply from the very next interaction. Scenarios like a tool-call notification

Reference-Free Evaluation of Taxonomies

ResearchDGX agent

arXiv:2505.11470v3 Announce Type: replace Abstract: We introduce two reference-free metrics for quality evaluation of taxonomies in the absence of labels. The first metric evaluates robustness by calc

SEEK: Steering LLM Reasoning for RAG via Internal Reasoning Sketches

ResearchDGX agent

arXiv:2601.09402v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by incorporating external knowledge into the generation process. Benefiti

SigmaScale: LLM Compression with SVD-based Low-Rank Decomposition and Learned Scaling Matrices

Model ReleasesDGX agent

arXiv:2606.07098v1 Announce Type: new Abstract: We present SigmaScale, a method for learning auxiliary scaling matrices S to aid truncated Singular Value Decomposition (SVD) based Large Language Model

Signal-Driven Observation for Long-Horizon Web Agents

AgentsDGX agent

arXiv:2606.06708v1 Announce Type: new Abstract: Web agents operating over long horizons ingest raw DOM and accessibility trees -- routinely tens of thousands of tokens -- at every action step, causing

Style or Content? Evaluating Style Classifiers with Controlled Content Overlap

Model ReleasesDGX agent

arXiv:2606.07103v1 Announce Type: new Abstract: Style classifiers can use content cues that correlate with style labels in naturally collected data, yet we lack a systematic way to measure this relian

SWE-Explore: Benchmarking How Coding Agents Explore Repositories

Model ReleasesDGX agent

arXiv:2606.07297v1 Announce Type: cross Abstract: Repository-level coding benchmarks such as SWE-bench have driven a rapid surge in the capabilities of coding agents. Yet they usually treat coding tas

Sycophantic Praise: Evaluating Excessive Praise in Language Models

SafetyDGX agent

arXiv:2606.07441v1 Announce Type: new Abstract: Sycophancy in language models is typically studied as excessive agreement or validation, while explicit praise and flattery have received comparatively

TA-RAG: Tone-Aware Retrieval-Augmented Generation for Peer-Support Health Communication

ResearchDGX agent

arXiv:2606.06794v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) successfully grounds large language model (LLM) outputs in trusted documents, but factual grounding alone is insuff

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling

SafetyDGX agent

arXiv:2507.06419v3 Announce Type: replace Abstract: Reward modeling (RM), which captures human preferences to align large language models (LLMs), is increasingly employed in tasks such as model finetu

The Dark Regulome: Disentangling Predictability from Regulation in Genomic Foundation Models

Local AiDGX agent

arXiv:2606.06834v1 Announce Type: new Abstract: High-grade gliomas integrate into neural circuits through functional synapses with neurons, raising the question of which noncoding elements shape synap

The Lipreading Gap: Do VSR Models Perceive Visual Speech Like Human Lipreaders?

ResearchDGX agent

arXiv:2606.07435v1 Announce Type: cross Abstract: Visual speech recognition (VSR) models now surpass human lipreaders on benchmarks, but do such gains establish human-like visual speech perception? To

The Necessity of Setting Temperature in LLM-as-a-Judge

ResearchDGX agent

arXiv:2603.28304v2 Announce Type: replace Abstract: Using large language models (LLMs) as judges for evaluating model outputs has emerged as an important paradigm for automated evaluation. However, th

The Piggyback Hypothesis of Generalization: Explaining and Mitigating Emergent Misalignment

Model ReleasesDGX agent

arXiv:2606.06667v1 Announce Type: new Abstract: The mechanisms behind LLMs' broad over-generalization beyond training examples remain unclear. Emergent misalignment (EM) offers a striking case study:

Translate-R1: Cost-Aware Translation Tool Use via Reinforcement Learning

SafetyDGX agent

arXiv:2606.06835v1 Announce Type: new Abstract: The performance gap across languages in LLMs is well documented, and closing it natively requires pretraining or fine-tuning on corpora that, for most l

Tree-of-Experience: A Structured Experience-Management Solution for Self-Evolving Agents under Low-Repetition and Implicit-Reward Environments

Model ReleasesDGX agent

arXiv:2606.06960v1 Announce Type: new Abstract: Experience-based self-evolution is crucial for LLM agents, but existing benchmarks often assume explicit goals, stable task patterns, and clear feedback

UnpredictaBench: A Benchmark for Evaluating Distributional Randomness in LLMs

Model ReleasesDGX agent

arXiv:2606.06622v1 Announce Type: new Abstract: We introduce UnpredictaBench, an evaluation that tests the ability of large language models (LLMs) to capture true underlying distributions. As LLMs are

What Do People Actually Want From AI? Mapping Preference Plurality

SafetyDGX agent

arXiv:2606.06674v1 Announce Type: new Abstract: Large Language Models (LLMs) are often fine-tuned through Reinforcement Learning from Human Feedback (RLHF) to align with people's preferences and value

When Better Codebooks Are Not Enough: Predictive Performance and Behavioral Reliability in LLM Political Event Coding

ResearchDGX agent

arXiv:2606.06781v1 Announce Type: new Abstract: High accuracy does not necessarily make an LLM a faithful coder. This issue matters because many social-science studies rely on expert-written codebooks

When to Think Deeply: Inhibitory Deliberation for LLM Reasoning

Model ReleasesDGX agent

arXiv:2606.06745v1 Announce Type: new Abstract: Reasoning Large Language Models can improve problem-solving performance through deliberative inference, but invoking slow reasoning for every input is c

Your UnEmbedding Matrix is Secretly a Feature Lens for Text Embeddings

ResearchDGX agent

arXiv:2606.07502v1 Announce Type: new Abstract: Large language models exhibit impressive zero-shot capabilities across a wide range of downstream tasks. However, they struggle to function as off-the-s

5 Jun 2026

A Komi-Yazva--Russian Parallel Corpus and Evaluation Protocol for Zero- and Few-Shot LLM Translation

SafetyDGX agent

arXiv:2606.06420v1 Announce Type: new Abstract: We present the first Komi-Yazva--Russian parallel corpus together with an explicit evaluation protocol for studying LLM translation in an endangered, ex

A Model of Multi-turn Human Persuadability Using Probabilistic Belief Tracing

ResearchDGX agent

arXiv:2606.05330v1 Announce Type: new Abstract: Large language models can shift human beliefs across high-stakes domains, but most persuasion studies rely on pre/post belief change. These endpoint mea

A Survey on Diffusion Language Models

ResearchDGX agent

arXiv:2508.10875v3 Announce Type: replace Abstract: Diffusion Language Models (DLMs) are rapidly emerging as a powerful and promising alternative to the dominant autoregressive (AR) paradigm. By gener

A Systematic Analysis of Biases in Large Language Models

SafetyDGX agent

arXiv:2512.15792v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have rapidly become indispensable tools for acquiring information and supporting human decision-making. However,

ABBEL: Learning Natural-Language Belief States for Memory-Efficient Interaction

AgentsDGX agent

arXiv:2512.20111v2 Announce Type: replace Abstract: As the time horizons of sequential decision-making tasks grow, keeping full interaction histories in model context becomes increasingly costly. Rece

ACE-SQL: Adaptive Co-Optimization via Empirical Credit Assignment for Text-to-SQL

SafetyDGX agent

arXiv:2606.05906v1 Announce Type: new Abstract: Text-to-SQL maps natural language questions to executable SQL queries. Modern databases often contain large and complex schemas, making schema linking a

Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM

ResearchDGX agent

arXiv:2510.05544v2 Announce Type: replace Abstract: Large language models (LLM) and vision-language models (VLM) have achieved state-of-the-art performance, but they impose significant memory and comp

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding

AgentsDGX agent

arXiv:2512.05774v2 Announce Type: replace-cross Abstract: Long video understanding (LVU) is challenging because answering real-world queries often depends on sparse, temporally dispersed cues buried i

AdaPlanBench: Evaluating Adaptive Planning in Large Language Model Agents under World and User Constraints

Model ReleasesDGX agent

arXiv:2606.05622v1 Announce Type: new Abstract: Planning for real-world problems by language models often involves both world and user constraints, which may not be fully specified upfront and are pro

AdaPLD: Adaptive Retrieval and Reuse for Efficient Model-Free Speculative Decoding

ResearchDGX agent

arXiv:2606.05742v1 Announce Type: new Abstract: Speculative decoding accelerates generation by verifying multiple drafted tokens in a single target-model forward pass, reducing sequential decoding ite

Agents' Last Exam

Model ReleasesDGX agent

arXiv:2606.05405v1 Announce Type: cross Abstract: Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deploym

Aligning Tree-Search Policies with Fixed Token Budgets in Test-Time Scaling of LLMs

Model ReleasesDGX agent

arXiv:2602.09574v2 Announce Type: replace Abstract: Tree-search decoding is an effective form of test-time scaling for large language models (LLMs), but real-world deployment often imposes a fixed per

Alignment Risks from Capability-Seeking RL Training

SafetyDGX agent

arXiv:2602.12124v2 Announce Type: replace-cross Abstract: While most AI alignment research focuses on preventing models from generating explicitly harmful content, a more subtle risk arises from capab

Almieyar-Oryx-BloomBench: A Bilingual Multimodal Benchmark for Cognitively Informed Evaluation of Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.05531v1 Announce Type: cross Abstract: Despite the rapid progress of Vision-Language Models (VLMs), the field lacks benchmarks that rigorously diagnose their true reasoning abilities and ch

An Embarrassingly Simple Detector for Model Extraction Attacks in Large Language Model API Traffic

Model ReleasesDGX agent

arXiv:2606.05725v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed through hosted APIs, making model extraction a practical threat to model ownership and service

An ERP Study on Recursive Locative Processing in Mandarin-Speaking Children with Autism

ResearchDGX agent

arXiv:2606.05620v1 Announce Type: new Abstract: Recursion enables the generation of hierarchical linguistic structures but imposes substantial processing demands during real-time comprehension. While

Analysis of the Neglect-Zero Effect in Large Language Models

SafetyDGX agent

arXiv:2606.05864v1 Announce Type: new Abstract: We investigate the extent to which the language processing of LLMs resembles human cognitive processes, focusing on a human cognitive bias called the ex

ArcANE: Do Role-Playing Language Agents Stay in Character at the Right Time?

Model ReleasesDGX agent

arXiv:2606.05553v1 Announce Type: new Abstract: Role-playing language agents (RPLAs) should play characters whose values and behavior evolve as the story progresses, not maintain a fixed persona. Exis

Ask or Assume? Uncertainty-Aware Clarification-Seeking in Coding Agents

AgentsDGX agent

arXiv:2603.26233v2 Announce Type: replace Abstract: As Large Language Model (LLM) agents are increasingly deployed in open-ended domains like software engineering, they frequently encounter underspeci

Asuka-Bench: Benchmarking Code Agents on Underspecified User Intent and Multi-Round Refinement

Model ReleasesDGX agent

arXiv:2606.05920v1 Announce Type: cross Abstract: Existing code-generation benchmarks score a single mapping from a complete prompt to a one-shot output. However, real web development is different. Us

AURA: Intent-Directed Probing for Implicit-Need Surfacing in Situated LLM Agents

Model ReleasesDGX agent

arXiv:2606.05557v1 Announce Type: new Abstract: A situated query like 'where is Lin Wei?' often encodes more than its literal content: the user may also want to know whether Lin Wei is free, in a good

Automatic Labelling of Speech Translation Errors

ResearchDGX agent

arXiv:2606.06047v1 Announce Type: new Abstract: Errors in speech translations reduce trustworthiness of Speech Translation (ST) systems and can have serious consequences. Yet currently there is no est

Benchmarking Open-Source Layout Detection Models for Data Snapshot Extraction from Institutional Documents

Model ReleasesDGX agent

arXiv:2606.06242v1 Announce Type: new Abstract: Institutional documents contain substantial amounts of operational and analytical information embedded within figures and tables. Current approaches for

Better Literary Translation: A Multi-Aspect Data Generation and LLM Training Approach

Model ReleasesDGX agent

arXiv:2606.05924v1 Announce Type: new Abstract: Literary translation poses unique challenges due to the scarcity of high-quality annotated data and the need to balance expression fluency with literary

Beyond Alignment: Value Diversity as a Collective Property in Multicultural Agent Systems

SafetyDGX agent

arXiv:2606.05985v1 Announce Type: new Abstract: Multicultural multi-agent systems are increasingly deployed in globally diverse settings, where different agents are grounded in different cultural back

Beyond tokens: a unified framework for latent communication in LLM-based multi-agent systems

SafetyDGX agent

arXiv:2606.05711v1 Announce Type: new Abstract: Multi-agent systems built on large language models (LLMs) have become a prevailing paradigm for tackling complex reasoning, planning, and tool-use tasks

Bootstrapping Semantic Layer from Execution for Text-to-SQL

ApplicationsDGX agent

arXiv:2606.05634v1 Announce Type: new Abstract: Real-world text-to-SQL is often under-specified until user phrases are grounded in how the database stores values. Prior work attempts to address this b

CaliDist: Calibrating Large Language Models via Behavioral Robustness to Distraction

ResearchDGX agent

arXiv:2606.05799v1 Announce Type: cross Abstract: Existing calibration methods for Large Language Models (LLMs) often overlook a critical dimension of trustworthiness: a model's {em behavioral robustn

Can LLMs Be Constrained to the Past? Improving Knowledge Cutoff through Recall-Based Prompting

Model ReleasesDGX agent

arXiv:2606.05804v1 Announce Type: new Abstract: Prompted knowledge cutoff instructs a large language model (LLM) to act as if information beyond a specified cutoff date were unavailable. However, prio

CHALIS: A Challenge Dataset for Language Identification in Difficult Scenarios

Model ReleasesDGX agent

arXiv:2606.06088v1 Announce Type: new Abstract: We present CHALIS (Challenging Language Identification Samples), a new benchmark dataset explicitly designed to address difficult cases in language iden

Channel-Wise Mixed-Precision Quantization for Large Language Models

Model ReleasesDGX agent

arXiv:2410.13056v4 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated remarkable success across a wide range of language tasks, but their deployment on edge devices remain

ChartAttack: Testing the Vulnerability of LLMs to Malicious Prompting in Chart Generation

ResearchDGX agent

arXiv:2601.12983v3 Announce Type: replace Abstract: Multimodal large language models (MLLMs) are increasingly used to automate chart generation from data tables, improving analysis and reporting effic

CHASE: Adversarial Red-Blue Teaming for Improving LLM Safety using Reinforcement Learning

SafetyDGX agent

arXiv:2606.05523v1 Announce Type: new Abstract: Despite advances in safety alignment, prompt-rewriting attacks such as persona modulation, fictional framing and persuasion-based reformulation, can byp

'Chi nas dal soch el sent de legn' -- Auditing Text Corpora for Lombard

SafetyDGX agent

arXiv:2606.06349v1 Announce Type: new Abstract: Several of the world's languages are still under-resourced in terms of Natural Language Processing (NLP) tools. This is mostly due to the lack of high-q

CLASH: Evaluating Language Models on Judging High-Stakes Dilemmas from Multiple Perspectives

Model ReleasesDGX agent

arXiv:2504.10823v4 Announce Type: replace Abstract: Navigating dilemmas involving conflicting values is challenging even for humans in high-stakes domains, let alone for AI, yet prior work has been li

CLFEC: A New Task for Unified Linguistic and Factual Error Correction in paragraph-level Chinese Professional Writing

Model ReleasesDGX agent

arXiv:2602.23845v2 Announce Type: replace Abstract: Chinese text correction has traditionally focused on spelling and grammar, while factual error correction is usually treated separately. However, in

← Previous
1…4041424344…129
Next →