AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
5 Jun 2026

Learning What to Forget: Improving LLM Unlearning via Learned Token-Level Importance

ResearchDGX agent

arXiv:2606.06320v1 Announce Type: cross Abstract: Machine unlearning aims to remove targeted knowledge from a trained model while preserving its general capabilities. For autoregressive language model

Less is MoE: Trimming Experts in Domain-Specialist Language Models

Model ReleasesDGX agent

arXiv:2606.05538v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models achieve strong performance through conditional computation, but their large parameter footprint poses deployment chall

Leveraging Large Language Models for Generating Research Topic Ontologies: A Multi-Disciplinary Study

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2508.20693v2 Announce Type: replace-cross Abstract: Ontologies and taxonomies of research fields are critical for managing and organising scientific knowledge, as they facilitate efficient class

LLM-Enhanced Dialogue Management for Full-Duplex Spoken Dialogue Systems

ResearchDGX agent

arXiv:2502.14145v3 Announce Type: replace Abstract: Achieving full-duplex communication in spoken dialogue systems (SDS) requires real-time coordination between listening, speaking, and thinking. This

LLMs Can Leak Training Data But Do They Want To? A Propensity-Aware Evaluation of Memorization in LLMs

ResearchDGX agent

arXiv:2606.06286v1 Announce Type: new Abstract: Large language models can reproduce training data, but existing memorization evaluations mostly measure whether models can be forced to do so, rather th

Localizing Prompt Ambiguity in Large Language Models with Probe-Targeted Attribution

Model ReleasesDGX agent

arXiv:2606.05486v1 Announce Type: new Abstract: Prompt ambiguity is a common source of failure in large language models, but is difficult to localize because it is a latent property of the prompt, whi

LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video

Model ReleasesDGX agent

arXiv:2606.05677v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have advanced image and video understanding and can increasingly handle longer visual inputs. Long-horizon ta

LoRi: Low-Rank Distillation for Implicit Reasoning

Model ReleasesDGX agent

arXiv:2606.05315v1 Announce Type: new Abstract: Implicit chain-of-thought (iCoT) methods aim to internalize reasoning in large language models, but often underperform explicit CoT prompting. We empiri

Many Circuits, One Mechanism: Input Variation and Evaluation Granularity in Circuit Discovery

ResearchDGX agent

arXiv:2606.06267v1 Announce Type: new Abstract: Circuit discovery methods identify subgraphs that explain specific model behaviors, and structural differences between discovered circuits are commonly

MARDoc: A Memory-Aware Refinement Agent Framework for Multimodal Long Document QA

AgentsDGX agent

arXiv:2606.05749v1 Announce Type: new Abstract: Iterative retrieval-reasoning agents have recently shown promise for multimodal long-document question answering. However, most existing systems maintai

MASF: A Multi-Model Adaptive Selection Framework for Abstractive Text summarization

ResearchDGX agent

arXiv:2606.05494v1 Announce Type: new Abstract: Automatic text summarization has become increasingly important due to the rapid growth of digital textual information. This paper presents a Multi-Model

MCBench: A Multicontext Safety Assessment Benchmark for Omni Large Language Models

Model ReleasesDGX agent

arXiv:2606.05177v1 Announce Type: new Abstract: Existing multimodal safety benchmarks focus solely on visual inputs and cannot assess Omni Large Language Models (LLMs) that process vision, audio, and

MDP-GRPO: Stabilized Group Relative Policy Optimization for Multi-Constraint Instruction Following

Model ReleasesDGX agent

arXiv:2606.06058v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards is ideal for multi-constraint instruction following, yet standard group-relative policy optimization (G

Measuring the sensitivity of LLM-based structured extraction to prompt, model, and schema choices in clinical discharge summaries

ResearchDGX agent

arXiv:2606.05970v1 Announce Type: new Abstract: Large language models are increasingly used for structured extraction from clinical free-text notes, but the sensitivity of their output to upstream con

Mechanistic Insights into Functional Sparsity in Multimodal LLMs via CoRe Heads

Local AiDGX agent

arXiv:2606.05843v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) demonstrate remarkable proficiency on complex vision-language tasks, the mechanisms by which they extract

Membrane: A Self-Evolving Contrastive Safety Memory for LLM Agent Defense

SafetyDGX agent

arXiv:2606.05743v1 Announce Type: cross Abstract: Despite advances in safety alignment, large language models remain vulnerable to continuously evolving jailbreaks. Existing fine-tuned safety classifi

MemoryCard: Topic-Aware Multi-Modal Clue Compression for Long-Video Question Answering

ResearchDGX agent

arXiv:2606.05917v1 Announce Type: cross Abstract: Long-video question answering remains challenging for Vision-Language Models (VLMs), as answer-relevant evidence is often sparse, transient, and tempo

MIRAI: Prediction and Generation of High-Impact Academic Research

ResearchDGX agent

arXiv:2606.05443v1 Announce Type: cross Abstract: The rapid pace of scientific publishing has made the identification and synthesis of high-impact work an increasingly urgent challenge. We introduce M

MLEvolve: A Self-Evolving Framework for Automated Machine Learning Algorithm Discovery

AgentsDGX agent

arXiv:2606.06473v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly applied to long-horizon tasks such as scientific discovery and machine learning engineering (MLE),

Multi-Granularity Reasoning for Natural Language Inference

ResearchDGX agent

arXiv:2606.05181v1 Announce Type: new Abstract: Natural Language Inference (NLI) is a fundamental task in natural language understanding that requires determining the logical relationship between a pr

Multi-task Learning is Not Enough: Representational Entanglement in Dual-output Second Language Speech Recognition

ResearchDGX agent

arXiv:2606.06065v1 Announce Type: new Abstract: Second-language (L2) speech recognition often requires transcriptions of pronunciations and intended meanings. Multi-task learning (MTL) is a natural ap

Multilingual Coreference Resolution via Cycle-Consistent Machine Translation

ResearchDGX agent

arXiv:2606.05444v1 Announce Type: new Abstract: Coreference resolution is a core NLP task, having a broad range of downstream applications, e.g.~machine translation, question answering, document summa

Multilingual Detection of Alzheimer's Disease from Speech: A Cross-Linguistic Transfer Learning Approach

ResearchDGX agent

arXiv:2606.05545v1 Announce Type: new Abstract: The development of multilingual Alzheimer's Disease Dementia (AD) detection models presents significant challenges due to the resource-intensive and tim

Narrative Knowledge Weaver: Narrative-Centric Retrieval-Augmented Reasoning for Long-Form Text Understanding

ResearchDGX agent

arXiv:2606.05724v1 Announce Type: new Abstract: Long-form narrative QA requires reasoning over evolving story worlds rather than isolated passages: answers may depend on earlier goals, changing charac

NAVIRA: Decoupled Stochastic Remasking for Masked Diffusion Language Models

Local AiDGX agent

arXiv:2606.06031v1 Announce Type: new Abstract: Masked diffusion language models generate text by iteratively unmasking many tokens in parallel, but this speed comes with a correction problem: tokens

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions

AgentsDGX agent

arXiv:2602.05843v2 Announce Type: replace Abstract: The rapid advancement of Large Language Models (LLMs) has catalyzed the development of autonomous agents capable of navigating complex environments.

On Advantage Estimates for Max@K Policy Gradients

SafetyDGX agent

arXiv:2606.06080v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards is widely used for post-training reasoning models, but sparse outcome rewards make exploration difficul

OneReason Technical Report

ApplicationsDGX agent

arXiv:2606.06260v1 Announce Type: cross Abstract: Generative recommendation models in the OneRec family have been widely deployed in many real-world services, such as short-video, live-streaming, adve

Operation-Guided Progressive Human-to-AI Text Transformation Benchmark for Multi-Granularity AI-Text Detection

Model ReleasesDGX agent

arXiv:2606.06481v1 Announce Type: new Abstract: As AI writing assistants become increasingly integrated into real-world drafting and revision workflows, many documents are no longer purely human-writt

OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation

SafetyDGX agent

arXiv:2606.06096v1 Announce Type: cross Abstract: Policy-gradient methods usually optimize expected return, but many real world applications care about distributional properties of returns: tail risk,

Ousiometrics: The essence of meaning aligns with a power-danger-structure framework instead of valence-arousal-dominance

SafetyDGX agent

arXiv:2110.06847v3 Announce Type: replace Abstract: From work emerging through the middle of the 20th century, the essence of meaning has become widely accepted as being described by the three orthogo

Ouvia: A User-centered Framework for Measuring Usability of Speech Translation in Real-World Communication Scenarios

ApplicationsDGX agent

arXiv:2606.06177v1 Announce Type: new Abstract: Speech translation (ST) is increasingly adopted in user applications, yet its evaluation largely focuses on decontextualized testbeds and holistic quali

PEFT of SLM for Telecommunications Customer Support: A Comparative Study of LoRA Configurations with Energy Consumption Analysis

Model ReleasesDGX agent

arXiv:2606.05176v1 Announce Type: new Abstract: While large language models (LLMs) show strong performance in natural language understanding and generation, their evaluation and adaptation to domain-s

Pitfalls of Evaluating Language Models with Open Benchmarks

SafetyDGX agent

arXiv:2507.00460v3 Announce Type: replace Abstract: Open Large Language Model (LLM) benchmarks, such as HELM and BIG-Bench, provide standardized and transparent evaluation protocols that support compa

PlanBench-V: A Spatial Planning Map Benchmark for Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.05744v1 Announce Type: new Abstract: Spatial planning maps are central to territorial governance, translating planning objectives, regulations, and spatial strategies into visual forms for

Predict and Reconstruct: Joint Objectives for Self-Supervised Language Representation Learning

Model ReleasesDGX agent

arXiv:2606.05173v1 Announce Type: new Abstract: Masked language modelling (MLM) has been the dominant pre-training objective for text encoders since BERT, yet it encourages representations that are st

Predictable Scaling Laws of Optimal Hyperparameters for LLM Continued Pre-training

ResearchDGX agent

arXiv:2606.05610v1 Announce Type: new Abstract: The efficacy of continued pre-training for Large Language Models (LLMs) hinges upon hyperparameter configurations, such as learning rate and batch size.

ProSarc: Prosody-Aware Sarcasm Recognition Framework via Temporal Prosodic Incongruity

Local AiDGX agent

arXiv:2606.06168v1 Announce Type: cross Abstract: We present ProSarc, an audio-only framework that detects sarcasm by modelling temporal prosodic incongruity, that is, the mismatch between local proso

ProSPy: A Profiling-Driven SQL-Python Agentic Framework for Enterprise Text-to-SQL

Model ReleasesDGX agent

arXiv:2606.05836v1 Announce Type: new Abstract: Large language models have substantially advanced Text-to-SQL systems, yet applying them to enterprise-scale databases remains challenging. Real-world d

QueryAgent-R1: Bridging Query Generation and Product Retrieval for E-Commerce Query Recommendation

SafetyDGX agent

arXiv:2606.05671v1 Announce Type: new Abstract: Query recommendation in e-commerce search aims to proactively suggest queries that match users' potential interests. However, existing methods mainly op

ReasoningFlow: Discourse Structures for Understanding LLM Reasoning Traces

Model ReleasesDGX agent

arXiv:2606.05402v1 Announce Type: new Abstract: Large reasoning models (LRMs) produce reasoning traces with non-linear structures, such as backtracking and self-correction, that complicate the evaluat

RedditPersona: A Modular Framework for Community-Conditioned LLM Adaptation from Reddit

Model ReleasesDGX agent

arXiv:2606.06027v1 Announce Type: cross Abstract: Community-conditioned language model adaptation requires choices about data collection, community definition, and evaluation that are currently made i

Reducing Hallucinations in Complex Question Answering using Simple Graph-based Retrieval-Augmented Generation (long version)

Model ReleasesDGX agent

arXiv:2606.05901v1 Announce Type: new Abstract: Large language models (LLMs) have fundamentally transformed the landscape of Natural Language Processing. Despite these advances, LLMs and LLM-based sys

Reinforcement Learning Elicits Contextual Learning of Unseen Language Translation

ResearchDGX agent

arXiv:2606.06428v1 Announce Type: new Abstract: Prior work has shown that large language models (LLMs) can translate unseen or low-resource languages by undergoing continued training or even by encodi

Representing Research Attention as Contextually Structured Flows

Model ReleasesDGX agent

arXiv:2606.05895v1 Announce Type: new Abstract: Research attention is widely used as an indicator of visibility, influence, and societal uptake, yet it is typically represented as aggregated counts th

Rethinking LoRA Memory Through the Lens of KV Cache Compression

Model ReleasesDGX agent

arXiv:2606.05698v1 Announce Type: new Abstract: Parametric retrieval augmentation encodes document information into lightweight, document-specific modules such as LoRA adapters, reducing the need to i

ReTreVal: Reasoning Tree with Validation and Cross-Problem Memory for Large Language Models

ResearchDGX agent

arXiv:2601.02880v2 Announce Type: replace-cross Abstract: Every existing inference-time reasoning framework discards all failure context at problem boundaries, leaving a model solving problem 500 no w

Retrospective Harness Optimization: Improving LLM Agents via Self-Preference over Trajectory Rollouts

AgentsDGX agent

arXiv:2606.05922v1 Announce Type: cross Abstract: AI agents rely on a harness of skills, tools, and workflows to solve complex problems. Continually improving this harness is essential for adapting to

ReverseEOL: Improving Training-free Text Embeddings via Text Reversal in Decoder-only LLMs

ResearchDGX agent

arXiv:2606.05858v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) have opened new avenues for generating training-free text embeddings. However, the causal attention in d

Revising Context, Shifting Simulated Stance: Auditing LLM-Based Stance Simulation in Online Discussions

ResearchDGX agent

arXiv:2606.06443v1 Announce Type: new Abstract: Large language models are increasingly used to simulate social media users and infer how individuals may respond to online discussions. However, it rema

Revisiting Lexicon Evaluation in Unsupervised Word Discovery

SafetyDGX agent

arXiv:2606.06183v1 Announce Type: cross Abstract: Building a lexicon from discovered word-like units is a central goal in zero-resource speech processing. But do our evaluations provide a trustworthy

Scaffold, Not Vocabulary? A Controlled, Two-Tier, Pre-Registered Study of a Popperian Code-Generation Skill

Model ReleasesDGX agent

arXiv:2606.06454v1 Announce Type: cross Abstract: Large language models increasingly write, review, and judge code, and a fast-growing practice equips them with prompt 'skills' that ask the model to r

Seeing is Believing? Evaluating Vision-Language Model Susceptibility in Agent-to-Agent Multimodal Persuasion

Model ReleasesDGX agent

arXiv:2510.22768v2 Announce Type: replace Abstract: As autonomous agents increasingly interact, they inevitably attempt to influence one another. While prior work in text-only settings has explored th

Self-Augmenting Retrieval for Diffusion Language Models

TutorialsDGX agent

arXiv:2606.06474v1 Announce Type: new Abstract: Discrete diffusion language models generate text by iteratively denoising an entire response in parallel. At each step, they predict tentative tokens fo

Self-supervised User Profile Generation for Personalization

Model ReleasesDGX agent

arXiv:2606.05336v1 Announce Type: new Abstract: Personalizing large language models (LLMs) has become a central challenge as LLMs are deployed across recommendation, search, dialogue, and content gene

Semi-Offline Reinforcement Learning for Optimized Text Generation

ResearchDGX agent

arXiv:2306.09712v2 Announce Type: replace-cross Abstract: In reinforcement learning (RL), there are two major settings for interacting with the environment: online and offline. Online methods explore

SkillComposer: Learning to Evolve Agent Skills for Specification and Generalization

AgentsDGX agent

arXiv:2606.06079v1 Announce Type: new Abstract: Agent skills, which consist of reusable strategies that guide agent reasoning and action, have shown strong potential for improving model capability at

SoCRATES: Towards Reliable Automated Evaluation of Proactive LLM Mediation across Domains and Socio-cognitive Variations

Model ReleasesDGX agent

arXiv:2606.05563v1 Announce Type: cross Abstract: Evaluating LLM mediators remains challenging, as mediation unfolds as a real-time trajectory shaped by disputants' shifting emotions, intentions, and

SpanNorm: Reconciling Training Stability and Performance in Deep Transformers

ResearchDGX agent

arXiv:2601.22580v2 Announce Type: replace Abstract: The success of Large Language Models (LLMs) hinges on the stable training of deep Transformer architectures. A critical design choice is the placeme

Stability vs. Manipulability: Evaluating Robustness Under Post-Decision Interaction in LLM Judges

Model ReleasesDGX agent

arXiv:2606.05384v1 Announce Type: cross Abstract: LLM-as-judge evaluation is widely used in benchmarking pipelines, where model outputs are compared and ranked using automated evaluators. These pipeli

← Previous
1…4243444546…129
Next →