AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
5 Jun 2026

Code2LoRA: Hypernetwork-Generated Adapters for Code Language Models under Software Evolution

Model ReleasesDGX agent

arXiv:2606.06492v1 Announce Type: cross Abstract: Code language models need repository-level context to resolve imports, APIs, and project conventions. Existing methods inject this knowledge as long i

Coding with 'Enemy': Can Human Developers Detect AI Agent Sabotage?

Model ReleasesDGX agent

arXiv:2606.05647v1 Announce Type: cross Abstract: AI coding agents are increasingly embedded in real-world software development, collaborating with human developers while gaining broader access to cod

ColBERTSaR: Sparsified ColBERT Index via Product Quantization

ResearchDGX agent

arXiv:2606.05568v1 Announce Type: cross Abstract: While ColBERT is an effective neural retrieval architecture, it requires a heavy index structure to support candidate set retrieval based on approxima


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

CollabBench: Benchmarking and Unleashing Collaborative Ability of LLMs with Diverse Players via Proactive Engagement

Model ReleasesDGX agent

arXiv:2606.05793v1 Announce Type: new Abstract: While LLM-based agents excel at individual tasks, effective collaboration with realistic human partners remains challenging. Most of the existing conver

CollabSim: A CSCW-Grounded Methodology for Investigating Collaborative Competence of LLM Agents through Controlled Multi-Agent Experiments

AgentsDGX agent

arXiv:2606.06399v1 Announce Type: new Abstract: Multi-agent systems (MAS) built on large language models have shown growing promise, with their effectiveness resting on agents' ability to coordinate t

CoMoL: Efficient Mixture of LoRA Experts via Dynamic Core Space Merging

Model ReleasesDGX agent

arXiv:2603.00573v2 Announce Type: replace Abstract: Large language models (LLMs) achieve remarkable performance on diverse downstream and domain-specific tasks via parameter-efficient fine-tuning (PEF

ComplexityMT: Benchmarking the Interaction Between Text Complexity and Machine Translation

ResearchDGX agent

arXiv:2606.05421v1 Announce Type: new Abstract: When a text is translated, does the translation retain the complexity of the original? We introduce ComplexityMT, a new challenge for assessing how text

Compress-Distill: Reasoning Trace Compression for Efficient Knowledge Distillation

Model ReleasesDGX agent

arXiv:2606.05988v1 Announce Type: cross Abstract: Reasoning models produce long chain-of-thought traces that are costly to distill and encourage verbose student outputs. We study post-hoc compression

Contextualized Prompting For Stance Detection On Social Media

Model ReleasesDGX agent

arXiv:2606.06022v1 Announce Type: new Abstract: Stance detection on social media is challenging due to short, noisy, and context-dependent language. While large language models (LLMs) show zero-shot g

Continual Learning Bench: Evaluating Frontier AI Systems in Real-World Stateful Environments

Model ReleasesDGX agent

arXiv:2606.05661v1 Announce Type: cross Abstract: Continual learning, the ability of AI systems to improve through sequential experience, has attracted substantial interest, but no high-quality benchm

Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering

ResearchDGX agent

arXiv:2510.05709v2 Announce Type: replace-cross Abstract: LLM benchmarking metrics often misstate performance and uncertainty as they rely on two assumptions that frequently do not hold in practice: (

CoT-Space: A Theoretical Framework for Internal Slow-Thinking via Reinforcement Learning

ResearchDGX agent

arXiv:2509.04027v3 Announce Type: replace-cross Abstract: Test-time scaling, primarily manifested through multi-step Chain-of-Thought (CoT) reasoning via Reinforcement Learning (RL), has emerged as a

Decomposing Factual Sycophancy in Language Models: How Size and Instruction Tuning Shape Robustness

ResearchDGX agent

arXiv:2606.06306v1 Announce Type: new Abstract: Factual sycophancy occurs when a language model abandons a correct, verifiable answer under social pressure. Because a flip occurs only when pressure to

Dense Contexts Are Hard Contexts: Lexical Density Limits Effective Context in LLMs

Model ReleasesDGX agent

arXiv:2606.06203v1 Announce Type: new Abstract: Input length and the position of relevant information are widely cited as the primary causes of degraded LLM long-context performance. Here, we study le

DiG-Plan: Mitigating Early Commitment for Tool-Graph Planning via Diffusion Guidance

ResearchDGX agent

arXiv:2606.05728v1 Announce Type: cross Abstract: Generating executable tool plans requires selecting appropriate subsets from tool libraries, a combinatorial search problem with an exponentially larg

Do MLLMs Capture How Interfaces Guide User Behavior? A Benchmark for Multimodal UI/UX Design Understanding

Model ReleasesDGX agent

arXiv:2505.05026v5 Announce Type: replace Abstract: User interface (UI) design goes beyond visuals to shape user experience (UX), underscoring the shift toward UI/UX as a unified concept. While recent

DocHop-QA: Towards Multi-Hop Reasoning over Multimodal Document Collections

Model ReleasesDGX agent

arXiv:2508.15851v2 Announce Type: replace Abstract: Despite rapid progress in large language models (LLMs), current QA benchmarks still overlook the core challenge of real-world scientific information

Domain-Aware Mispronunciation Detection and Diagnosis Using Language-Specific Statistical Graphs

Model ReleasesDGX agent

arXiv:2606.05569v1 Announce Type: new Abstract: Mispronunciation Detection and Diagnosis (MDD) has gained increasing importance in computer-assisted language learning and speech technology in recent y

Domain-Conditioned Safety in Frontier Computer-Using Agents: A 793-Episode Browser Benchmark, a Coding-Domain Cross-Reference, and a Reproducibility Audit of Recent Red-Teaming

Model ReleasesDGX agent

arXiv:2606.05233v1 Announce Type: cross Abstract: Recent computer-using-agent (CUA) red-teaming papers report prompt-injection attack success rates (ASR) of 42-98%, but these headline numbers cluster

Dynamic Thinking-Token Selection for Efficient Reasoning in Large Reasoning Models

ResearchDGX agent

arXiv:2601.18383v2 Announce Type: replace-cross Abstract: Large Reasoning Models (LRMs) excel at solving complex problems by explicitly generating a reasoning trace before deriving the final answer. H

EDIT: Evidence-Diagnosed Intervention Training for Rule-Faithful LLM Grading

ApplicationsDGX agent

arXiv:2606.06350v1 Announce Type: new Abstract: Reliable rubric grading requires more than accurate score prediction. Each judgement must be grounded in the mark scheme and evidence from the student a

Efficient Punctuation Restoration via Weighted Lookahead Scoring Method for Streaming ASR Systems

Model ReleasesDGX agent

arXiv:2606.05179v1 Announce Type: new Abstract: Punctuation restoration improves ASR (Automatic Speech Recognition) readability. However streaming ASR requires online decisions with limited future con

EGTR-Review: Efficient Evidence-Grounded Scientific Peer Review Generation via Multi-Agent Teacher Distillation

AgentsDGX agent

arXiv:2606.06025v1 Announce Type: new Abstract: Scientific peer review generation has attracted increasing attention for reducing reviewing burdens and providing timely feedback. However, existing Lar

EMBER: Efficient Memory via Budgeted Evidence Retention for Long-Horizon Agents

SafetyDGX agent

arXiv:2606.05894v1 Announce Type: new Abstract: Long-horizon agents can archive large histories, but future answers still incur retrieval, rereading, and context costs. When retained memory misses ans

Emergent Language as an Approach to Conscious AI

AgentsDGX agent

arXiv:2606.06380v1 Announce Type: new Abstract: The question of whether artificial systems can be conscious remains open, in part because existing approaches either evaluate systems against theory-der

English-to-Prakrit Machine Translation via Multilingual Transfer Learning

Model ReleasesDGX agent

arXiv:2606.06038v1 Announce Type: new Abstract: We study English-to-Prakrit machine translation in a low-resource setting where the target language is unsupported by IndicTrans2. We adapt the multilin

Epidemiology of Model Collapse: Modeling Synthetic Data Contamination via Bilayer SIR Dynamics

Model ReleasesDGX agent

arXiv:2606.05168v1 Announce Type: new Abstract: Training on synthetic data causes model collapse, but existing analyses treat this as single-chain degradation. In reality, the AI ecosystem involves cr

EpiEvolve: Self-Evolving Agents for Streaming Pandemic Forecasting under Regime Shifts

AgentsDGX agent

arXiv:2606.05513v1 Announce Type: cross Abstract: Epidemic LLM forecasters are usually trained and evaluated as static supervised models, whereas operational pandemic forecasting is a streaming proces

Epistemic Injustice in Language Models: An Audit of Pretraining Filters and Guardrails

SafetyDGX agent

arXiv:2606.05936v1 Announce Type: new Abstract: Modern language models rely on pretraining filters to remove undesirable content from training corpora and inference-time guardrails to suppress undesir

Evaluating Stochastic Collapse and Implicit Bias in Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2606.05874v1 Announce Type: new Abstract: Current evaluations for Multimodal Large Language Models (MLLMs) overwhelmingly focus on utility-driven objectives, leaving model behavior under logic-n

Executable Schema Contracts: From Automatic Ingestion to Multi-Source Retrieval

AgentsDGX agent

arXiv:2606.05415v1 Announce Type: new Abstract: Real-world data spans tables, documents, and semi-structured files with implicit semantics. Querying this data requires integrating evidence across inco

Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations

Local AiDGX agent

arXiv:2510.17256v2 Announce Type: replace Abstract: Large language models have exhibited impressive performance across a broad range of downstream tasks in natural language processing. However, how a

FiLM-Based Speaker Conditioning of a SpeechLLM for Pathological Speech Recognition

Model ReleasesDGX agent

arXiv:2606.06211v1 Announce Type: new Abstract: Automatic speech recognition (ASR) has advanced remarkably for standard speech; however, pathological speech from neurological conditions remains a sign

Forgive or forget: Understanding the context of hate in audio retrieval systems

SafetyDGX agent

arXiv:2606.05857v1 Announce Type: new Abstract: Handling toxic retrieval in text-to-audio systems is challenging due to contextual dependencies. Existing strategies (e.g., rephrasing, summarization) r

FOXGLOVE: Understanding Goal-Oriented and Anchored Writing Feedback from Experts and LLMs on Argumentative Essays

ResearchDGX agent

arXiv:2606.06271v1 Announce Type: new Abstract: While large language models (LLMs) are increasingly used to generate writing feedback, there remains no systematic comparison of LLM and expert feedback

Framing, Judging, Steering: An Assessable Competency Model for Teach-ing Students to Reason With Generative AI

ResearchDGX agent

arXiv:2606.05983v1 Announce Type: cross Abstract: Generative AI makes answers easy and understanding hard, and uncritical use invites cognitive offloading. Schools still measure unaided performance, y

From Scoring to Explanations: Evaluating SHAP and LLM Rationales for Rubric-based Teaching Quality Assessment

ResearchDGX agent

arXiv:2606.05180v1 Announce Type: new Abstract: Automated scoring models are increasingly used to assign rubric-based quality ratings to complex language performances, including classroom transcripts,

From Self to Other: Evaluating Demographic Perspective-Taking in LLM Hate Speech Annotation

Model ReleasesDGX agent

arXiv:2606.06266v1 Announce Type: new Abstract: Hate speech detection is inherently subjective: people from different demographic groups perceive the same content very differently. Collecting enough a

Generic Triple-Latent Compression with Gated Associative Retrieval

Model ReleasesDGX agent

arXiv:2606.05175v1 Announce Type: new Abstract: We study generic triple-latent sequence models that maintain a running token state and compressed pair-memory pathway to capture higher-order token inte

GLASS: GRPO-Trained LoRA for Acoustic Style Steering in Zero-Shot Text-to-Speech

SafetyDGX agent

arXiv:2606.05889v1 Announce Type: cross Abstract: We propose GLASS, a framework for composable acoustic style control in zero-shot autoregressive text-to-speech (TTS) that learns controls from post-ge

Grounded but Misleading: Evaluating Semantic Alignment in AI-Generated Security Explanations

SafetyDGX agent

arXiv:2602.05056v2 Announce Type: replace-cross Abstract: Online scams increasingly leverage fluent and context-aware social engineering strategies, creating growing demand for AI systems that explain

Harnessing Generalist Agents for Contextualized Time Series

AgentsDGX agent

arXiv:2606.05404v1 Announce Type: cross Abstract: Time series are often embedded in rich contexts that are essential for holistic modeling. Moreover, real-world practitioners often require end-to-end

Harnessing Structural Context for Entity Alignment Foundation Models

Model ReleasesDGX agent

arXiv:2606.06109v1 Announce Type: new Abstract: Entity alignment (EA) aims to identify equivalent entities across heterogeneous knowledge graphs (KGs) and is a key component of knowledge fusion and cr

Human Adults and LLMs as Scientists: Who Benefits from Active Exploration?

SafetyDGX agent

arXiv:2606.06464v1 Announce Type: new Abstract: A long-standing finding in the causal learning literature is that adults struggle to identify conjunctive causal rules, where an effect requires the sim

Humans' ALMANAC: A Human Collaboration Dataset of Action-Level Mental Model Annotations for Agent Collaboration

Model ReleasesDGX agent

arXiv:2606.06388v1 Announce Type: cross Abstract: Recent advances in LLM agents have enabled complex cognitive capabilities, such as multi-step reasoning, planning, and tool use, that increasingly pos

IA-RAG: Interval-Algebra-Driven Temporal Reasoning for Dynamic Knowledge Retrieval

Model ReleasesDGX agent

arXiv:2606.06044v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) has shown strong effectiveness in grounding Large Language Models (LLMs) with external knowledge. However, existing

IDEAL: Leveraging Infinite and Dynamic Characterizations of Large Language Models for Query-focused Summarization

SafetyDGX agent

arXiv:2407.10486v3 Announce Type: replace-cross Abstract: Query-focused summarization (QFS) aims to produce summaries that answer particular questions of interest, enabling greater user control and pe

Improving Answer Extraction in Context-based Question Answering Systems Using LLMs

Model ReleasesDGX agent

arXiv:2606.06197v1 Announce Type: new Abstract: Question answering (QA) systems have achieved notable progress with the advent of large language models (LLMs). However, they still face challenges in a

Improving Heart-Focused Medical Question Answering in LLMs via Variance-Aware Rubric Rewards with GRPO

Local AiDGX agent

arXiv:2606.05174v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown strong promise in healthcare applications. Yet deploying general-purpose models in real-world settings remains d

InfoDensity: Rewarding Information-Dense Traces for Efficient Reasoning

ResearchDGX agent

arXiv:2603.17310v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) with extended reasoning capabilities often generate verbose and redundant reasoning traces, incurring unnecessary

InfoShield: Privacy-Preserving Speech Representations for Mental Health Screening via Information-Theoretic Optimization

ResearchDGX agent

arXiv:2606.05561v1 Announce Type: new Abstract: Speech-based mental health screening offers scalable depression detection, yet clinical deployment faces a significant barrier: users' privacy concerns

Interpreting Style Representations via Style-Eliciting Prompts

ResearchDGX agent

arXiv:2606.05716v1 Announce Type: new Abstract: Style representation learning is a powerful tool for authorship analysis and modeling writing style, yet the latent nature of learned representations ma

IR3DE: A Linear Router for Large Language Models

ResearchDGX agent

arXiv:2606.06098v1 Announce Type: new Abstract: Foundational Large Language Models (LLMs) demonstrate proficiency on a wide range of general tasks, and achieve remarkable results on various specialize

LANTERN: Layered Archival and Temporal Episodic Retrieval Network for Long-Context LLM Conversations

ApplicationsDGX agent

arXiv:2606.05182v1 Announce Type: new Abstract: Large language models discard critical details when conversation history is compacted to fit within finite context windows. We present LANTERN (Layered

Large Language Models are Perplexed by some Political Parties

SafetyDGX agent

arXiv:2606.05937v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used, including in political applications, but their political fairness has been little studied. We assess

Latent Reasoning with Normalizing Flows

SafetyDGX agent

arXiv:2606.06447v1 Announce Type: new Abstract: Large language models often improve reasoning by generating explicit chain-of-thought (CoT), demonstrating the importance of intermediate computation. H

LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents

Model ReleasesDGX agent

arXiv:2606.06087v1 Announce Type: new Abstract: Agent systems increasingly use textual skills to encode reusable task procedures, but injecting these skills into the prompt at every step incurs substa

LeanMarathon: Toward Reliable AI Co-Mathematicians through Long-Horizon Lean Autoformalization

Local AiDGX agent

arXiv:2606.05400v1 Announce Type: cross Abstract: Long-horizon autoformalization of research mathematics fails not only at hard lemmas, but at scale: statements drift, dependencies tangle, context dec

Learning Self-Correction in Vision-Language Models via Rollout Augmentation

TutorialsDGX agent

arXiv:2602.08503v2 Announce Type: replace-cross Abstract: Self-correction is essential for solving complex reasoning problems in vision-language models (VLMs). However, existing reinforcement learning

Learning to Route LLMs from Implicit Cost-Performance Preferences via Meta-Learning

ResearchDGX agent

arXiv:2606.06178v1 Announce Type: cross Abstract: Large language models (LLMs) present a trade-off between performance and cost, where more powerful models incur greater expense. LLM routing aims to m

← Previous
1…4142434445…129
Next →