AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
Research

MAGE: All-[MASK] Block Already Knows Where to Look in Block Diffusion LLM

DGX agent

arXiv:2602.14209v2 Announce Type: replace-cross Abstract: Block diffusion LLMs are an emerging paradigm for parallel language generation, but their KV caching makes memory access the dominant bottlene

researcharxiv-cs-cl
8 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

Meaning in Order, Order in Meaning: Semantic R-precision for Keyphrase Evaluation

DGX agent

arXiv:2606.07057v1 Announce Type: cross Abstract: Evaluating the quality of automatically generated keyphrases remains a complex challenge. Traditional metrics either rely on exact lexical matching or

researcharxiv-cs-cl
8 Jun 2026
Tutorials

Mechanistic Evidence for Faithfulness Decay in Chain-of-Thought Reasoning

DGX agent

arXiv:2602.11201v2 Announce Type: replace Abstract: Chain-of-Thought (CoT) explanations are widely used to interpret how language models solve complex problems, yet it remains unclear whether these st

tutorialsarxiv-cs-cl
8 Jun 2026
Safety

Mining Useful General Data for Low-Resource Domain Adaptation

DGX agent

arXiv:2511.07380v2 Announce Type: replace Abstract: Adapting large language models (LLMs) to low-resource domains remains challenging due to the scarcity of domain-specific data. While in-domain data

safetyarxiv-cs-cl
8 Jun 2026
Model Releases

MMAE: A Massive Multitask Audio Editing Benchmark

DGX agent

arXiv:2606.07229v1 Announce Type: cross Abstract: We introduce MMAE, a Massive Multitask Audio Editing benchmark, serving as the first comprehensive evaluation testbed designed for general-purpose ins

model-releasesarxiv-cs-cl
8 Jun 2026
Model Releases

mmPISA-bench: Do LLMs Reason Equally Well Across 43 Languages?

DGX agent

arXiv:2606.07069v1 Announce Type: new Abstract: We introduce mmPISA-bench, a compact high-quality multilingual reasoning benchmark derived from the OECD Programme for International Student Assessment

model-releasesarxiv-cs-cl
8 Jun 2026
Research

Modeling semantic association in self-paced reading with language model embeddings

DGX agent

arXiv:2606.07066v1 Announce Type: new Abstract: Semantic association between a word and its context has been identified as an important component of reading comprehension, even when word predictabilit

researcharxiv-cs-cl
8 Jun 2026
Research

Modular Monolingual Adaptation using Pretrained Language Models

DGX agent

arXiv:2606.06738v1 Announce Type: new Abstract: Building monolingual language models (LMs) for low-resource languages typically relies on adapting pretrained language models (PLMs) by finetuning the w

researcharxiv-cs-cl
8 Jun 2026
Research

Multiscale POD of Transformer Attention Fields: Scale-Selective Analysis via Morlet Scalogram

DGX agent

arXiv:2606.06573v1 Announce Type: cross Abstract: We introduce scale-selective Proper Orthogonal Decomposition (POD) for transformer attention fields, inspired by the use of POD for extracting energet

researcharxiv-cs-cl
8 Jun 2026
Model Releases

Phun-Bench: Evaluating LLMs on Phonological Understanding in Chinese

DGX agent

arXiv:2606.07300v1 Announce Type: new Abstract: Language is a vehicle for thought, intricately tied to sounds, symbols, and meaning. However, most large language model (LLM) research focuses on meanin

model-releasesarxiv-cs-cl
8 Jun 2026
Research

PolarQuant: Leveraging Polar Transformation for Efficient Key Cache Quantization and Decoding Acceleration

DGX agent

arXiv:2502.00527v2 Announce Type: replace-cross Abstract: The KV cache in large language models is a dominant factor in memory usage, limiting their broader applicability. Quantizing the cache to lowe

researcharxiv-cs-cl
8 Jun 2026
Model Releases

Principles of Concept Representation in Sentence Encoders

DGX agent

arXiv:2606.06994v1 Announce Type: new Abstract: What makes a sentence encoder produce good concept representations? We approach this through the lens of representational compositionality: an encoder s

model-releasesarxiv-cs-cl
8 Jun 2026
Model Releases

PromptPrint: Behavioral Biometrics Through Natural Language Prompting in LLMs

DGX agent

arXiv:2606.06755v1 Announce Type: new Abstract: Authorship attribution research has traditionally focused on long-form, expressive texts; however, interactions with large language models (LLMs) are ty

model-releasesarxiv-cs-cl
8 Jun 2026
Research

Quantifying Media Representation Dynamics Across 25 Years of News Reporting on Policing-related Deaths

DGX agent

arXiv:2606.06812v1 Announce Type: new Abstract: We perform the largest known computational analysis of Canadian news narratives about police-involved deaths, spanning 4,000 articles from the last quar

researcharxiv-cs-cl
8 Jun 2026
Safety

RASFT: Rollout-Adaptive Supervised Fine-Tuning for Reasoning

DGX agent

arXiv:2606.07006v1 Announce Type: cross Abstract: Supervised fine-tuning (SFT) is a prevailing method for adapting large language models to reasoning tasks by imitating offline expert demonstrations,

safetyarxiv-cs-cl
8 Jun 2026
Model Releases

RECAP: Regression Evaluation for Continual Adaptation of Prompts

DGX agent

arXiv:2606.06698v1 Announce Type: cross Abstract: Production agentic systems routinely face evolving constraints and must comply from the very next interaction. Scenarios like a tool-call notification

model-releasesarxiv-cs-cl
8 Jun 2026
Research

Reference-Free Evaluation of Taxonomies

DGX agent

arXiv:2505.11470v3 Announce Type: replace Abstract: We introduce two reference-free metrics for quality evaluation of taxonomies in the absence of labels. The first metric evaluates robustness by calc

researcharxiv-cs-cl
8 Jun 2026
Research

SEEK: Steering LLM Reasoning for RAG via Internal Reasoning Sketches

DGX agent

arXiv:2601.09402v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by incorporating external knowledge into the generation process. Benefiti

researcharxiv-cs-cl
8 Jun 2026
Model Releases

SigmaScale: LLM Compression with SVD-based Low-Rank Decomposition and Learned Scaling Matrices

DGX agent

arXiv:2606.07098v1 Announce Type: new Abstract: We present SigmaScale, a method for learning auxiliary scaling matrices S to aid truncated Singular Value Decomposition (SVD) based Large Language Model

model-releasesarxiv-cs-cl
8 Jun 2026
Agents

Signal-Driven Observation for Long-Horizon Web Agents

DGX agent

arXiv:2606.06708v1 Announce Type: new Abstract: Web agents operating over long horizons ingest raw DOM and accessibility trees -- routinely tens of thousands of tokens -- at every action step, causing

agentsarxiv-cs-cl
8 Jun 2026
Model Releases

Style or Content? Evaluating Style Classifiers with Controlled Content Overlap

DGX agent

arXiv:2606.07103v1 Announce Type: new Abstract: Style classifiers can use content cues that correlate with style labels in naturally collected data, yet we lack a systematic way to measure this relian

model-releasesarxiv-cs-cl
8 Jun 2026
Model Releases

SWE-Explore: Benchmarking How Coding Agents Explore Repositories

DGX agent

arXiv:2606.07297v1 Announce Type: cross Abstract: Repository-level coding benchmarks such as SWE-bench have driven a rapid surge in the capabilities of coding agents. Yet they usually treat coding tas

model-releasesarxiv-cs-cl
8 Jun 2026
Safety

Sycophantic Praise: Evaluating Excessive Praise in Language Models

DGX agent

arXiv:2606.07441v1 Announce Type: new Abstract: Sycophancy in language models is typically studied as excessive agreement or validation, while explicit praise and flattery have received comparatively

safetyarxiv-cs-cl
8 Jun 2026
Research

TA-RAG: Tone-Aware Retrieval-Augmented Generation for Peer-Support Health Communication

DGX agent

arXiv:2606.06794v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) successfully grounds large language model (LLM) outputs in trusted documents, but factual grounding alone is insuff

researcharxiv-cs-cl
8 Jun 2026
Safety

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling

DGX agent

arXiv:2507.06419v3 Announce Type: replace Abstract: Reward modeling (RM), which captures human preferences to align large language models (LLMs), is increasingly employed in tasks such as model finetu

safetyarxiv-cs-cl
8 Jun 2026
Local Ai

The Dark Regulome: Disentangling Predictability from Regulation in Genomic Foundation Models

DGX agent

arXiv:2606.06834v1 Announce Type: new Abstract: High-grade gliomas integrate into neural circuits through functional synapses with neurons, raising the question of which noncoding elements shape synap

local-aiarxiv-cs-cl
8 Jun 2026
Research

The Lipreading Gap: Do VSR Models Perceive Visual Speech Like Human Lipreaders?

DGX agent

arXiv:2606.07435v1 Announce Type: cross Abstract: Visual speech recognition (VSR) models now surpass human lipreaders on benchmarks, but do such gains establish human-like visual speech perception? To

researcharxiv-cs-cl
8 Jun 2026
Research

The Necessity of Setting Temperature in LLM-as-a-Judge

DGX agent

arXiv:2603.28304v2 Announce Type: replace Abstract: Using large language models (LLMs) as judges for evaluating model outputs has emerged as an important paradigm for automated evaluation. However, th

researcharxiv-cs-cl
8 Jun 2026
Model Releases

The Piggyback Hypothesis of Generalization: Explaining and Mitigating Emergent Misalignment

DGX agent

arXiv:2606.06667v1 Announce Type: new Abstract: The mechanisms behind LLMs' broad over-generalization beyond training examples remain unclear. Emergent misalignment (EM) offers a striking case study:

model-releasesarxiv-cs-cl
8 Jun 2026
Safety

Translate-R1: Cost-Aware Translation Tool Use via Reinforcement Learning

DGX agent

arXiv:2606.06835v1 Announce Type: new Abstract: The performance gap across languages in LLMs is well documented, and closing it natively requires pretraining or fine-tuning on corpora that, for most l

safetyarxiv-cs-cl
8 Jun 2026
Model Releases

Tree-of-Experience: A Structured Experience-Management Solution for Self-Evolving Agents under Low-Repetition and Implicit-Reward Environments

DGX agent

arXiv:2606.06960v1 Announce Type: new Abstract: Experience-based self-evolution is crucial for LLM agents, but existing benchmarks often assume explicit goals, stable task patterns, and clear feedback

model-releasesarxiv-cs-cl
8 Jun 2026
Model Releases

UnpredictaBench: A Benchmark for Evaluating Distributional Randomness in LLMs

DGX agent

arXiv:2606.06622v1 Announce Type: new Abstract: We introduce UnpredictaBench, an evaluation that tests the ability of large language models (LLMs) to capture true underlying distributions. As LLMs are

model-releasesarxiv-cs-cl
8 Jun 2026
Safety

What Do People Actually Want From AI? Mapping Preference Plurality

DGX agent

arXiv:2606.06674v1 Announce Type: new Abstract: Large Language Models (LLMs) are often fine-tuned through Reinforcement Learning from Human Feedback (RLHF) to align with people's preferences and value

safetyarxiv-cs-cl
8 Jun 2026
Research

When Better Codebooks Are Not Enough: Predictive Performance and Behavioral Reliability in LLM Political Event Coding

DGX agent

arXiv:2606.06781v1 Announce Type: new Abstract: High accuracy does not necessarily make an LLM a faithful coder. This issue matters because many social-science studies rely on expert-written codebooks

researcharxiv-cs-cl
8 Jun 2026
Model Releases

When to Think Deeply: Inhibitory Deliberation for LLM Reasoning

DGX agent

arXiv:2606.06745v1 Announce Type: new Abstract: Reasoning Large Language Models can improve problem-solving performance through deliberative inference, but invoking slow reasoning for every input is c

model-releasesarxiv-cs-cl
8 Jun 2026
Research

Your UnEmbedding Matrix is Secretly a Feature Lens for Text Embeddings

DGX agent

arXiv:2606.07502v1 Announce Type: new Abstract: Large language models exhibit impressive zero-shot capabilities across a wide range of downstream tasks. However, they struggle to function as off-the-s

researcharxiv-cs-cl
8 Jun 2026
Safety

A Komi-Yazva--Russian Parallel Corpus and Evaluation Protocol for Zero- and Few-Shot LLM Translation

DGX agent

arXiv:2606.06420v1 Announce Type: new Abstract: We present the first Komi-Yazva--Russian parallel corpus together with an explicit evaluation protocol for studying LLM translation in an endangered, ex

safetyarxiv-cs-cl
5 Jun 2026
Research

A Model of Multi-turn Human Persuadability Using Probabilistic Belief Tracing

DGX agent

arXiv:2606.05330v1 Announce Type: new Abstract: Large language models can shift human beliefs across high-stakes domains, but most persuasion studies rely on pre/post belief change. These endpoint mea

researcharxiv-cs-cl
5 Jun 2026
Research

A Survey on Diffusion Language Models

DGX agent

arXiv:2508.10875v3 Announce Type: replace Abstract: Diffusion Language Models (DLMs) are rapidly emerging as a powerful and promising alternative to the dominant autoregressive (AR) paradigm. By gener

researcharxiv-cs-cl
5 Jun 2026
Safety

A Systematic Analysis of Biases in Large Language Models

DGX agent

arXiv:2512.15792v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have rapidly become indispensable tools for acquiring information and supporting human decision-making. However,

safetyarxiv-cs-cl
5 Jun 2026
Agents

ABBEL: Learning Natural-Language Belief States for Memory-Efficient Interaction

DGX agent

arXiv:2512.20111v2 Announce Type: replace Abstract: As the time horizons of sequential decision-making tasks grow, keeping full interaction histories in model context becomes increasingly costly. Rece

agentsarxiv-cs-cl
5 Jun 2026
Safety

ACE-SQL: Adaptive Co-Optimization via Empirical Credit Assignment for Text-to-SQL

DGX agent

arXiv:2606.05906v1 Announce Type: new Abstract: Text-to-SQL maps natural language questions to executable SQL queries. Modern databases often contain large and complex schemas, making schema linking a

safetyarxiv-cs-cl
5 Jun 2026
Research

Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM

DGX agent

arXiv:2510.05544v2 Announce Type: replace Abstract: Large language models (LLM) and vision-language models (VLM) have achieved state-of-the-art performance, but they impose significant memory and comp

researcharxiv-cs-cl
5 Jun 2026
Agents

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding

DGX agent

arXiv:2512.05774v2 Announce Type: replace-cross Abstract: Long video understanding (LVU) is challenging because answering real-world queries often depends on sparse, temporally dispersed cues buried i

agentsarxiv-cs-cl
5 Jun 2026
Model Releases

AdaPlanBench: Evaluating Adaptive Planning in Large Language Model Agents under World and User Constraints

DGX agent

arXiv:2606.05622v1 Announce Type: new Abstract: Planning for real-world problems by language models often involves both world and user constraints, which may not be fully specified upfront and are pro

model-releasesarxiv-cs-cl
5 Jun 2026
Research

AdaPLD: Adaptive Retrieval and Reuse for Efficient Model-Free Speculative Decoding

DGX agent

arXiv:2606.05742v1 Announce Type: new Abstract: Speculative decoding accelerates generation by verifying multiple drafted tokens in a single target-model forward pass, reducing sequential decoding ite

researcharxiv-cs-cl
5 Jun 2026
Model Releases

Agents' Last Exam

DGX agent

arXiv:2606.05405v1 Announce Type: cross Abstract: Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deploym

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

Aligning Tree-Search Policies with Fixed Token Budgets in Test-Time Scaling of LLMs

DGX agent

arXiv:2602.09574v2 Announce Type: replace Abstract: Tree-search decoding is an effective form of test-time scaling for large language models (LLMs), but real-world deployment often imposes a fixed per

model-releasesarxiv-cs-cl
5 Jun 2026
← Previous
1…5051525354…161
Next →