AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
49+ results
Syntheses

Synthesis: Arxiv-Cs-Cl

DGX agent

Auto-generated synthesis of 505 entries about arxiv-cs-cl

synthesisarxiv-cs-clauto-generated
16 Apr 2026
Safety

Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents

DGX agent

arXiv:2608.11110v1 Announce Type: new Abstract: When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compa

AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
safetyarxiv-cs-cl
12 Aug 2026
Research

ASR-Roundtrip Evaluation Can Mask Context- and Convention-Dependent Reading Errors in Chinese News TTS

DGX agent

arXiv:2608.10606v1 Announce Type: new Abstract: ASR-roundtrip evaluation is widely used as a scalable proxy for text-to-speech (TTS) intelligibility, but it can produce false negatives for reading err

researcharxiv-cs-cl
12 Aug 2026
Applications

Assessing Reliability of BERT-Based Models on Question Answering Tasks

DGX agent

arXiv:2608.10806v1 Announce Type: new Abstract: Reliability estimation of large language models is in many cases as crucial as their accuracy, as reliable models are more trustworthy, robust, and suit

applicationsarxiv-cs-cl
12 Aug 2026
Safety

Automated Data Enrichment using Confidence-Aware Fine-Grained Debate among Open-Source LLMs for Mental Health and Online Safety

DGX agent

arXiv:2512.06227v3 Announce Type: replace Abstract: Real-world indicators play an important role in many Natural Language Processing (NLP) applications, such as life events for mental health analysis

safetyarxiv-cs-cl
12 Aug 2026
Model Releases

Bayesian-Agent: Posterior-Guided Skill Evolution Across LLM Agent Harnesses

DGX agent

arXiv:2606.08348v2 Announce Type: replace Abstract: LLM agents increasingly rely on prompts, tools, memory, SOPs, skills, and harness feedback, yet current self-evolution pipelines often update these

model-releasesarxiv-cs-cl
12 Aug 2026
Research

Beyond a Bag of Features: Set-Level Instability in Sparse Autoencoders

DGX agent

arXiv:2608.11197v1 Announce Type: cross Abstract: Shani et al. (2026) show that LLM representations broadly recover human category boundaries, while failing to reflect fine-grained typicality structur

researcharxiv-cs-cl
12 Aug 2026
Research

Calibrating Post-Training Feature Shifts for LLM Data Contamination Detection

DGX agent

arXiv:2608.10462v1 Announce Type: new Abstract: Large language models (LLMs) are trained on massive and largely undisclosed corpora that may contain copyrighted or privacy-sensitive content. Data cont

researcharxiv-cs-cl
12 Aug 2026
Model Releases

Can Released LLM Vocabularies Support Token-Level Estimation of Hidden Corpora?

DGX agent

arXiv:2608.10690v1 Announce Type: new Abstract: Pretraining corpus composition shapes LLM capabilities, but it often remains hidden even when model weights are released. Prior work has inferred corpus

model-releasesarxiv-cs-cl
12 Aug 2026
Safety

Carefully Considering Culture: Analyzing LLM Alignment in Single- and Multi-Cultural Settings using Cultural Consensus Theory

DGX agent

arXiv:2608.09937v1 Announce Type: new Abstract: Recent work in NLP has probed large language models for their understanding of cultural norms across countries. However, this work typically considers d

safetyarxiv-cs-cl
12 Aug 2026
Local Ai

Certify or Refuse: A Cross-Model Map for Selective Risk Control with Coverage Floors under Covariate Shift

DGX agent

arXiv:2608.10893v1 Announce Type: new Abstract: Certified selective predictors attain whatever coverage they attain; operators impose an automation floor: answer at least a eta-fraction of shifted tar

local-aiarxiv-cs-cl
12 Aug 2026
Agents

Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design

DGX agent

arXiv:2608.10299v1 Announce Type: new Abstract: Agentic systems are increasingly expected to improve after deployment, yet single-entity self-evolution is often bounded by a static learning context, s

agentsarxiv-cs-cl
12 Aug 2026
Research

Conflict or Strategy? Asymmetric Role Framing of La France insoumise and Rassemblement National in French News Headlines, 2022-2025

DGX agent

arXiv:2608.09936v1 Announce Type: new Abstract: Do French news headlines frame left- and right-populist challengers as symmetric ``extremes,'' or as fundamentally different political adversaries? We e

researcharxiv-cs-cl
12 Aug 2026
Safety

ConRub-Med: Reinforcement Learning with Consensus Rubrics for Open-Ended Medical Question Answering

DGX agent

arXiv:2608.10996v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards has been especially effective in mathematics and coding, where answers can be checked automatically. Many

safetyarxiv-cs-cl
12 Aug 2026
Model Releases

Cost-Efficient Estimation of General Abilities Across Benchmarks

DGX agent

arXiv:2604.01418v2 Announce Type: replace Abstract: Thousands of diverse benchmarks have been developed to measure the quality of large language models (LLMs). Yet prior work has demonstrated that LLM

model-releasesarxiv-cs-cl
12 Aug 2026
Model Releases

Cracks in the Foundation: Seemingly Minor Architectural Choices Impact Long Context Extension

DGX agent

arXiv:2608.10296v1 Announce Type: new Abstract: One might imagine that architectural variations within the dense transformer paradigm have a limited effect on accuracy. However, we demonstrate that th

model-releasesarxiv-cs-cl
12 Aug 2026
Safety

Data Attribution of Emergent Misalignment with Persona Features

DGX agent

arXiv:2608.11025v1 Announce Type: new Abstract: Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A leadi

safetyarxiv-cs-cl
12 Aug 2026
Research

Decomposition-Induced Context-Memory Conflict: When Fact-Checking Pipelines Contradict Their Own Source Text

DGX agent

arXiv:2608.10627v1 Announce Type: new Abstract: Decompose-then-verify pipelines, including FActScore-style fact-checkers and long-form factuality evaluators, first split a passage into atomic claims b

researcharxiv-cs-cl
12 Aug 2026
Safety

Detecting an Effect Is Not Learning to Act on It: A Reward-SNR Floor for LLM Acquisition Agents

DGX agent

arXiv:2608.10441v1 Announce Type: cross Abstract: Many pipelines can pay a per-example cost to acquire an auxiliary, model-derived observation -- an LLM's structured reasoning, a slow oracle, an expen

safetyarxiv-cs-cl
12 Aug 2026
Model Releases

DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation

DGX agent

arXiv:2608.10636v1 Announce Type: cross Abstract: Visual document retrieval (VDR) is dominated by multi-billion-parameter models that are slow to index at full corpus scale and expensive to serve. Pri

model-releasesarxiv-cs-cl
12 Aug 2026
Model Releases

Do LLM Recommenders Know When They're Hallucinating? Auditing Confidence Calibration in Catalog Faithfulness

DGX agent

arXiv:2608.10008v1 Announce Type: cross Abstract: LLM recommenders for top-K item suggestion regularly emit titles outside the target catalog. Prior audits measure this as a binary out-of-domain rate;

model-releasesarxiv-cs-cl
12 Aug 2026
Safety

Dual-Loop Self-Evolution via Verifiable Emotion Feedback for Multi-Turn Empathetic Dialogue

DGX agent

arXiv:2608.10626v1 Announce Type: new Abstract: Large language models have demonstrated conversational capabilities, yet empathetic competence remains challenging. Empathetic support is inherently mul

safetyarxiv-cs-cl
12 Aug 2026
Safety

Enhancing Automated Essay Scoring With Three Techniques: Two-Stage Fine-Tuning, Score Alignment, and Self-Training

DGX agent

arXiv:2602.01747v2 Announce Type: replace Abstract: Automated Essay Scoring (AES) plays a crucial role in education by providing scalable and efficient assessment tools. However, in real-world setting

safetyarxiv-cs-cl
12 Aug 2026
Safety

Every Token Counts: Exact Likert-Scale Distributions for Measuring LLM Attitudes and Biases

DGX agent

arXiv:2608.10503v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly deployed as autonomous agents, accurately evaluating their latent values and biases is critical. The NL

safetyarxiv-cs-cl
12 Aug 2026
Research

EVIL-Detect for NLPCC 2026 Shared Task 6: LLM-Generated Text Detection

DGX agent

arXiv:2608.10698v1 Announce Type: new Abstract: The rapid development of large language models (LLMs) has increased the need for reliable detection of LLM-generated text, especially in realistic Chine

researcharxiv-cs-cl
12 Aug 2026
Research

Faster Superword Tokenization

DGX agent

arXiv:2604.05192v2 Announce Type: replace Abstract: Byte Pair Encoding (BPE) is a widely used tokenization algorithm, whose tokens cannot extend across pre-tokenization boundaries, functionally limiti

researcharxiv-cs-cl
12 Aug 2026
Safety

FoR-SALE: Frame of Reference-guided Spatial Adjustment in LLM-based Diffusion Editing

DGX agent

arXiv:2509.23452v2 Announce Type: replace-cross Abstract: Current text-to-image generation models, even state-of-the-art models, exhibit a significant performance gap when spatial expressions are desc

safetyarxiv-cs-cl
12 Aug 2026
Safety

Generation-Step-Aware Framework for Cross-Modal Representation and Control in Multilingual Speech-Text Models

DGX agent

arXiv:2601.17387v3 Announce Type: replace Abstract: Multilingual speech-text models rely on cross-modal language alignment to transfer knowledge between speech and text, but it remains unclear whether

safetyarxiv-cs-cl
12 Aug 2026
Model Releases

How Robust Are LLMs to Vietnamese Dialects?

DGX agent

arXiv:2608.10414v1 Announce Type: new Abstract: Large Language Models (LLMs) are typically evaluated on standard written Vietnamese, yet everyday communication frequently involves regional dialects th

model-releasesarxiv-cs-cl
12 Aug 2026
Model Releases

InSight-doc: Agentic Visual Perception for Long-Document Understanding

DGX agent

arXiv:2608.10628v1 Announce Type: cross Abstract: Long-document understanding often requires reasoning over many visually rich pages, making inference costly and prone to context rot. In this work, we

model-releasesarxiv-cs-cl
12 Aug 2026
Research

Leveraging Human Reading Behavior for Keyphrase Extraction: A Webcam-based Eye-tracking Corpus

DGX agent

arXiv:2608.10688v1 Announce Type: new Abstract: Purpose: Keyphrases are statistically and semantically important textual units that can also attract readers' attention during comprehension. However, e

researcharxiv-cs-cl
12 Aug 2026
Model Releases

Mapping and Measuring the Behavioral Evolution of Large Language Models

DGX agent

arXiv:2608.11027v1 Announce Type: cross Abstract: Benchmark leaderboards summarize how well a language model performs, but not how its behavior relates to that of other models or changes across genera

model-releasesarxiv-cs-cl
12 Aug 2026
Agents

Mitigating Context Interference for Reliable and Efficient Search Agents

DGX agent

arXiv:2608.10743v1 Announce Type: new Abstract: Recent research empowers Large Language Models (LLMs) as multi-turn search agents to iteratively retrieve and generate outputs until complex tasks are s

agentsarxiv-cs-cl
12 Aug 2026
Agents

MT-PingEval: Evaluating Multi-Turn Collaboration with Private Information Games

DGX agent

arXiv:2602.24188v2 Announce Type: replace Abstract: We present a scalable and verifiable methodology for evaluating language models in multi-turn interactions, using a suite of collaborative games tha

agentsarxiv-cs-cl
12 Aug 2026
Research

Multilingual Embedding Probes Fail to Generalize Across Learner Corpora

DGX agent

arXiv:2604.07095v2 Announce Type: replace Abstract: Do multilingual embedding models encode a language-general representation of proficiency? We investigate this by training linear and non-linear prob

researcharxiv-cs-cl
12 Aug 2026
Local Ai

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment

DGX agent

arXiv:2608.11167v1 Announce Type: cross Abstract: Existing Multimodal Large Language Models (MLLMs) predominantly rely on image-text pairs for modality alignment pretraining, mapping global image repr

local-aiarxiv-cs-cl
12 Aug 2026
Research

MUSE: A Full-Text Cross-Domain Knowledge Base of Scientific Problems, Solutions, and Rationales

DGX agent

arXiv:2608.10974v1 Announce Type: new Abstract: Scientific papers contain fine-grained records of problem solving: authors mention technical obstacles and methods that were used to address them, often

researcharxiv-cs-cl
12 Aug 2026
Model Releases

myMediWhisper: Construction of Burmese Medical Speech Corpus and Whisper Fine-Tuning for Clinical Dialogue ASR

DGX agent

arXiv:2608.11036v1 Announce Type: new Abstract: Although Whisper models benefit from large-scale multilingual pre-training, their performance on Burmese medical speech remains limited. This work prese

model-releasesarxiv-cs-cl
12 Aug 2026
Model Releases

No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding

DGX agent

arXiv:2503.05061v3 Announce Type: replace Abstract: Reliable evaluation of large language models (LLMs) is critical as their deployment rapidly expands, particularly in high-stakes domains such as bus

model-releasesarxiv-cs-cl
12 Aug 2026
Research

Off-Axis, On Purpose: Where a Transformer Computes Concepts and Why it Does So

DGX agent

arXiv:2608.10251v1 Announce Type: new Abstract: A transformer's answer lives on one axis: the direction its unembedding reads. Its intermediate states largely do not, and that off-axis position is usu

researcharxiv-cs-cl
12 Aug 2026
Agents

OpenPM: Auditable Point-in-Time Evaluation for LLM Portfolio-Management Agents

DGX agent

arXiv:2608.09988v1 Announce Type: cross Abstract: Large language models are increasingly used to read markets, assess risk, and allocate capital. However, reported results for LLM trading agents can b

agentsarxiv-cs-cl
12 Aug 2026
Research

Order Matters: LVLMs as Judges for Temporal Reasoning in Image Sequences

DGX agent

arXiv:2608.10908v1 Announce Type: cross Abstract: As generative multimedia evolves from static image synthesis to complex, interleaved visual narratives, a foundational bottleneck has emerged: the jud

researcharxiv-cs-cl
12 Aug 2026
Research

PERCEPT: A Corpus for POS Tagging and Analysis of Persian-English Code-Mixing

DGX agent

arXiv:2608.10109v1 Announce Type: new Abstract: Social media has become a major venue for multilingual communication, where users frequently mix multiple languages within a single utterance. Although

researcharxiv-cs-cl
12 Aug 2026
Safety

Position Encoding in Transformers: From Absolute and Relative Methods to Rotary Position Embeddings and Long-Context Scaling

DGX agent

arXiv:2608.10021v1 Announce Type: new Abstract: Self-attention models content-dependent interactions between tokens but does not by itself encode token order. Position encoding addresses this limitati

safetyarxiv-cs-cl
12 Aug 2026
Model Releases

Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference

DGX agent

arXiv:2608.10288v1 Announce Type: cross Abstract: The Large Language Model from Power Law Decoder Representations (PLDR-LLM) and its attention, Power Law Graph Attention (PLGA), replace the fixed bili

model-releasesarxiv-cs-cl
12 Aug 2026
Model Releases

REAP: Relation-Aware Elicitation and Parsing for Closed-Book Knowledge Base Construction from LLMs

DGX agent

arXiv:2608.10963v1 Announce Type: new Abstract: We present the REAP system for the AKBC Shared Task 2026 on constructing knowledge bases from language models in a closed-book setting, subject to a bud

model-releasesarxiv-cs-cl
12 Aug 2026
Research

Reinforcement Learning-based Semi-supervised Knowledge Distillation with LLM-as-a-Judge

DGX agent

arXiv:2604.02621v2 Announce Type: replace Abstract: Reinforcement Learning (RL) substantially improves the reasoning capabilities of language models, but most existing RL fine-tuning approaches rely e

researcharxiv-cs-cl
12 Aug 2026
Model Releases

ReRound: Reconstructive Rounding to Resolve Midpoint Ambiguity in Calibration-Free LLM Quantization

DGX agent

arXiv:2608.11045v1 Announce Type: cross Abstract: ReRound (Reconstructive Rounding) is a post-training quantization method that addresses the midpoint ambiguity inherent in standard round-to-nearest (

model-releasesarxiv-cs-cl
12 Aug 2026
← Previous
1
Next →
7,646 results
← Previous
123…160
Next →