AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
4 Aug 2026

MedTextWeaver: Procedural Knowledge Evolution in Agentic Medical Text Editing

AgentsDGX agent

arXiv:2602.00740v2 Announce Type: replace Abstract: Medical text editing is essential for improving communication among diverse stakeholders in clinical settings. However, adapting LLM agents to this

MedUPS: Towards Diagnostic Assistance in Uncommon Medical Cases with Large Language Models

SafetyDGX agent

arXiv:2608.01012v1 Announce Type: new Abstract: Uncommon and off-guideline cases are difficult for clinical decision support, because physicians must make a series of management decisions under diagno

MemoryForge: Synthesize Lifelong Memory for Human-Like LLM Agents

AgentsDGX agent

arXiv:2608.00007v1 Announce Type: new Abstract: Equipping Large Language Models (LLMs) with human-like personas is crucial for agentic applications, such as role-play and user simulation. Traditional


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

MemSIF: From Structured Interactions to Dual-Track Fact Memory for LLM Agents

ResearchDGX agent

arXiv:2608.01742v1 Announce Type: cross Abstract: Long-term memory is critical for LLM agents operating over long-horizon interactions. However, several persistent limitations of existing memory syste

Mind the Gap: Zero-Query Jailbreaks via Filter-Generator Discrepancy in Text-to-Image Systems

SafetyDGX agent

arXiv:2608.00973v1 Announce Type: new Abstract: Text-to-image (T2I) systems typically have prompt-level safety filters before the generator to block unsafe requests, yet such systems remain vulnerable

Morphology Aware Reversible Semantic Tokenization and Hierarchical Word Composition for Tamil Language Models

Model ReleasesDGX agent

arXiv:2608.01153v1 Announce Type: new Abstract: Statistical subword tokenizers can process arbitrary text, but their units need not align with lexical or grammatical structure. This is especially impo

Native Multilingual Chain-of-Thought Reasoning in Low-Resource Southeast Asian Languages

SafetyDGX agent

arXiv:2608.00533v1 Announce Type: new Abstract: Large Language Models have achieved substantial progress in reasoning capabilities. Yet in low-resource native settings, many suffer from cross-lingual

Neural Circuit Function Inference with LLMs

ResearchDGX agent

arXiv:2608.00059v1 Announce Type: new Abstract: The success of connectome mapping now shifts the challenge of understanding the nervous system to the interpretation of neural circuits. Here, we devise

No One Wins in Nuclear War: A Social Simulation of Military Decision-making

SafetyDGX agent

arXiv:2608.01868v1 Announce Type: cross Abstract: WOPR is a social-simulation environment for studying how organizations make high-stakes decisions, built on a deterministic, replay-validated rules en

Not the Dimension, the Norm: What Matters in Gradient-Free Weight Perturbation of Language Models

Model ReleasesDGX agent

arXiv:2608.01624v1 Announce Type: new Abstract: Adapting a language model to a task no longer requires training all of its weights, and a line of parameter-efficient methods has driven the trainable c

Observatorio Lazaro: A self-populating database of anglicism usage in the Spanish press

ResearchDGX agent

arXiv:2608.00713v1 Announce Type: new Abstract: This paper describes Observatorio Lazaro, a language resource that monitors unassimilated lexical borrowings (predominantly English lexical borrowings o

Obshazard-bench: Benchmarking Multimodal Foundation Models for Real-Time Disaster Intelligence from Raw Earth Observation Streams

Model ReleasesDGX agent

arXiv:2608.00012v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) are increasingly used to interpret Earth observation data, yet their capability to support real-world disaster

On the Wings of Imagination: Conflicting Script-based Multi-role Framework for Humor Caption Generation

ResearchDGX agent

arXiv:2602.06423v2 Announce Type: replace Abstract: Humor is a commonly used and intricate human language in daily life. Humor generation, especially in multi-modal scenarios, is a challenging task fo

OoO-Spec: Out-of-Order Semantic Speculation for Fast Tool Calling

Model ReleasesDGX agent

arXiv:2608.00814v1 Announce Type: new Abstract: LLMs generate tool calls token by token, even though the function choice and argument values can often be predicted in parallel from the request and too

OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution

Model ReleasesDGX agent

arXiv:2608.00677v1 Announce Type: new Abstract: AI agents operate in persistent environments where early state changes can influence decisions far into the future. Unlike conventional language-model i

OpenDebateEvidence: A Massive-Scale Argument Mining and Summarization Dataset

ResearchDGX agent

arXiv:2406.14657v4 Announce Type: replace Abstract: We introduce OpenDebateEvidence, a comprehensive dataset for argument mining and summarization sourced from the American Competitive Debate communit

Opt.Gear Technical Report

Model ReleasesDGX agent

arXiv:2608.01034v1 Announce Type: new Abstract: We introduce Opt.Gear, a foundation model designed for efficient on-device deployment, real-tim inference, and strong task capability. It includes a den

Orchestrating Dual-Boundaries: An Arithmetic Intensity Inspired Acceleration Framework for Diffusion Language Models

ResearchDGX agent

arXiv:2511.21759v2 Announce Type: replace Abstract: Diffusion-based large language models (dLLMs) have recently gained significant attention for their exceptional performance and inherent potential fo

OTAP: Structure-Aware Optimal Transport for Evaluating Planning and Execution in Agent Trajectories

AgentsDGX agent

arXiv:2607.17082v2 Announce Type: replace-cross Abstract: Large language model agents solve tasks by generating trajectories that interleave planning, tool calls, and intermediate results. Current eva

PALMs: Using Multi Construct-Grounded Rationales for Modeling Population Preferences in LLMs

SafetyDGX agent

arXiv:2608.01458v1 Announce Type: new Abstract: Large language models are being extensively used to simulate individual user behavior, yet faithfully representing a population requires capturing the s

Passing Coarse Marginal Checks Can Be Cheap: Persona Mixtures and Imprecise Treatment-Response Estimates in an LLM Persona Panel

Model ReleasesDGX agent

arXiv:2608.00979v1 Announce Type: cross Abstract: Large language models are increasingly used as synthetic research participants and are often validated by whether their marginal responses resemble hu

PGMem: Tightly Coupled Persona-Memory Graph for Lifelong Personalized Agents

AgentsDGX agent

arXiv:2608.01708v1 Announce Type: new Abstract: Long-term personalized dialogue agents must track user preferences as their personas evolve. Existing memory systems organize past events well, but stor

PICTURE: Enhancing Theory-of-Mind in Large Language Models by Revealing, Not Hiding, Characters' Lack of Knowledge

ResearchDGX agent

arXiv:2608.01598v1 Announce Type: new Abstract: Simulating human-like Theory of Mind (ToM) has been a longstanding problem in natural language processing (NLP). To address this, existing works introdu

PlainMedScale: A Corpus of Multi-Level Simplified Medical Texts in German and English

ResearchDGX agent

arXiv:2608.01158v1 Announce Type: new Abstract: We introduce PlainMedScale, a topic-aligned medical corpus spanning four levels of comprehensibility in German and English, drawn from MSD (professional

Practical Online KV Cache Compaction for LLM Agents: An Empirical Study

AgentsDGX agent

arXiv:2608.00902v1 Announce Type: new Abstract: LLM agents accumulate long trajectories of reasoning steps, tool calls, and environment feedback, making the KV cache a major inference bottleneck. KV c

PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise

Model ReleasesDGX agent

arXiv:2608.02372v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in task-oriented dialogue systems that support multi-step decision-making in high-stakes domains

Predicting Startup Exit from Textual Descriptors - A Computational Linguistics Framework

ResearchDGX agent

arXiv:2608.00045v1 Announce Type: new Abstract: This study shows that textual descriptors alone can predict early-stage startup success, defined as Exit, without relying on contextual, financial, or h

Progressive Agent Skill Generation via Reinforcement Learning

AgentsDGX agent

arXiv:2608.01678v1 Announce Type: cross Abstract: Existing skill generation methods largely rely on heuristics or pipeline-style consolidation, which must be specially designed for different evidence

Prompt-Induced Waste in Large Reasoning Models: A Preregistered Two-Harness Benchmark of Coding Agents

Model ReleasesDGX agent

arXiv:2608.01347v1 Announce Type: new Abstract: Large reasoning models used as coding agents incur costs from deliberation, tool calls, and repeated agent turns, yet the causal effect of prompt wordin

Pruned BPE: Post-training Visibility Pruning and Token Reallocation for Byte Pair Encoding

ResearchDGX agent

arXiv:2608.00837v1 Announce Type: new Abstract: Byte Pair Encoding (BPE) is widely used for subword tokenization, but standard BPE exposes every learned merge token to the downstream model, including

QR-Erase: Efficient Subspace-Based Machine Unlearning with Layer Localization

ResearchDGX agent

arXiv:2608.01422v1 Announce Type: new Abstract: Machine unlearning seeks to remove targeted information from trained models without requiring costly retraining. Existing optimization-based methods oft

Question Begets Question: Self-Evolving Curriculum for Reinforcement Fine-Tuning on Competition Mathematics

ResearchDGX agent

arXiv:2608.01522v1 Announce Type: cross Abstract: Teaching a language model a skill it has not mastered is obstructed by three recurring difficulties: training data is scarce, ground-truth reasoning t

Quick on the Uptake: Eliciting Implicit Intents from Human Demonstrations for Personalized Mobile-Use Agents

SafetyDGX agent

arXiv:2508.08645v3 Announce Type: replace Abstract: As multimodal large language models advance rapidly, the automation of mobile tasks has become increasingly feasible through the use of mobile-use a

Qwen-CUA: Native Computer Use for (almost) Everything

Model ReleasesDGX agent

arXiv:2608.02352v1 Announce Type: cross Abstract: Native computer use offers a general interface for agents to operate almost any software available to people, but requires long-horizon state tracking

RADAR: Rubric-Aware Dependency and Redundancy Analysis for LLM-as-Judge Evaluation

Model ReleasesDGX agent

arXiv:2608.01810v1 Announce Type: new Abstract: Rubric-based LLM-as-judge pipelines often assume that evaluation criteria provide independent signals. In practice, however, criteria can be behaviorall

RAG Strategies for Natural Language-Based SQL Query and REST API Call Generation

ApplicationsDGX agent

arXiv:2602.07086v2 Announce Type: replace-cross Abstract: Enterprise software systems commonly expose business functionality through both relational databases and REST APIs. Accessing these interfaces

RAGOCR: Optical Compression of Retrieval-Augmented Text via Visual Representation

ResearchDGX agent

arXiv:2608.00765v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) has become essential for knowledge-intensive question answering, yet scaling RAG pipelines remains challenging due

REFLEX: Rethinking MoE Inference as Refinement-Aware Compute Allocation in Diffusion Language Models

Model ReleasesDGX agent

arXiv:2608.01784v1 Announce Type: cross Abstract: Mixture-of-experts (MoE) models increase parameter capacity by activating only a small subset of experts for each token. This conditional-computation

RestoreKV: Recovering Full-Cache Behavior Under Aggressive Query-Agnostic KV Cache Eviction

Model ReleasesDGX agent

arXiv:2608.01247v1 Announce Type: new Abstract: Query-agnostic KV cache eviction compresses a context once and reuses the resulting cache for arbitrary future queries, but performance can collapse und

Rethinking and formalising the state across languages: a unified computational learning theory account

ResearchDGX agent

arXiv:2608.00523v1 Announce Type: new Abstract: The linguistic notion of state has traditionally been restricted to the construct (annexation) state of Afroasiatic languages and treated as a language-

Retrieval Augmented Biomedical Question Answering with Weak Question Recovery and Neural Reranking for BioASQ Task 14b

ResearchDGX agent

arXiv:2608.01468v1 Announce Type: new Abstract: This work presents DS@GT ARC BioASQ team's work for a biomedical question answering pipeline, integrating multi-source query expansion, neural reranking

Revisiting Generalization Across Difficulty Levels: It's Not So Easy

ResearchDGX agent

arXiv:2511.21692v2 Announce Type: replace Abstract: We investigate how well large language models (LLMs) generalize across different task difficulties, a key question for effective data curation and e

RH-RAG: Trustworthy Long-Form Generation for Privacy-Constrained Settings

Local AiDGX agent

arXiv:2608.01311v1 Announce Type: new Abstract: Generating long-form content from extensive internal reports remains challenging for organizations operating under strict privacy and security constrain

Rhythm of the Deep: Two-Tier Combinatorial Structure in Sperm Whale Codas Revealed by Acoustic Unit Induction

ResearchDGX agent

arXiv:2606.16084v2 Announce Type: replace-cross Abstract: Sperm-whale codas are conventionally described as recurring click-count and timing patterns. We show instead that their waveforms contain a tw

Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning on Frontier Science Benchmarks

Model ReleasesDGX agent

arXiv:2608.02442v1 Announce Type: cross Abstract: Scientific reasoning benchmarks typically evaluate large language models (LLMs) using final-answer accuracy. However, a correct answer does not necess

RING: Retrieval-Internalized Generation for Continual Large-Scale Knowledge Injection

Model ReleasesDGX agent

arXiv:2608.01630v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) improves factuality but adds latency and engineering overhead at serving time. We propose RING (Retrieval-Internali

RMSWeb: Reflection, Failure-Mode Mining, and Salvage-DS for Web Agent Reinforcement Learning

SafetyDGX agent

arXiv:2608.00335v1 Announce Type: cross Abstract: Compact web agents can reduce deployment cost, but training them poses challenges in both data collection and post-SFT reinforcement learning (RL). Su

Role Steering of Language Models for Social Simulations

Model ReleasesDGX agent

arXiv:2608.00023v1 Announce Type: new Abstract: Social simulations built from language-model agents need role-conditioned behavior that can be checked before agents are placed into a simulated populat

Romanized Arabic Across Dialects: Views, Usage Patterns, and Linguistic Variation

SafetyDGX agent

arXiv:2608.02555v1 Announce Type: new Abstract: Arabizi refers to Arabic written in Latin script. Although previous studies have shown that the prevalence and usage of Arabizi vary by factors such as

RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States

AgentsDGX agent

arXiv:2608.02508v1 Announce Type: cross Abstract: Learning-based memory systems for self-evolving LLM agents face two tightly coupled challenges. First, trajectory-indexed utilities grow with the inte

RubricReviewer: From Direct Critique to Objective and Comprehensive Rubric-Driven Peer Review

AgentsDGX agent

arXiv:2608.00005v1 Announce Type: new Abstract: Peer review at major venues is under unprecedented submission pressure, motivating the use of large language models (LLMs) as review assistants. Existin

S^4R: Selective Sampling, Subspaces, and Sparse Reconstruction for Compressed Long-Context KV Caching

Model ReleasesDGX agent

arXiv:2608.00528v1 Announce Type: new Abstract: The growth of context window lengths in Large Language Models (LLMs) significantly enhances their long-context capabilities but incurs prohibitive memor

Same violence, different answer: how AI responds to coercive control against women across languages

ResearchDGX agent

arXiv:2608.01436v1 Announce Type: cross Abstract: Women experiencing coercive control, a form of intimate partner violence increasingly conducted through digital devices, are turning to conversational

SCHEDBench: A Benchmark for Evaluating LLM Constraint Faithfulness in Natural-Language Combinatorial Scheduling

Model ReleasesDGX agent

arXiv:2608.00991v1 Announce Type: cross Abstract: This paper introduces SCHEDBench, a natural-language benchmark for evaluating combinatorial scheduling constraint faithfulness under surface-form vari

ScrambleToolBench: Agents Search Exhaustively Even When Their Own Map Points to the Next Step

Model ReleasesDGX agent

arXiv:2608.02358v1 Announce Type: new Abstract: To operate robustly in open-world environments, autonomous agents should be able to infer the behavior of unfamiliar systems through interaction alone,

SeDeM: Selective Decompression of Hidden-State Memories for Long-Context Question Answering

ResearchDGX agent

arXiv:2608.00311v1 Announce Type: new Abstract: Long-context inference with large language models (LLMs) is costly: self-attention during prefill scales quadratically with sequence length, and the key

Select-And-Extract: A Lightweight Plugin for Retrieval-Augmented Generation

ResearchDGX agent

arXiv:2608.00658v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) for language model (LM) systems fundamentally has two failure modes: retrieval failure and reading failure. The for

Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models

SafetyDGX agent

arXiv:2507.02778v3 Announce Type: replace Abstract: Although large language models (LLMs) have transformed AI, they still make errors and follow unproductive reasoning paths. Self-correction is vital

Self-Improving Large Language Models via Progressive Experience Evolution

SafetyDGX agent

arXiv:2608.02139v1 Announce Type: new Abstract: Large language models (LLMs) capable of self-improvement require not only effective policy optimization, but also a principled mechanism for transformin

Semantic Alignment of AI Models: Concept Collapse, Checkpoint Dynamics, and Cross-Lingual Transfer

SafetyDGX agent

arXiv:2608.01585v1 Announce Type: new Abstract: Language model benchmarking is a difficult task. Outcome reasoning alone does not test the model's conceptualization of language and popular open-source

← Previous
1…910111213…128
Next →