AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
30 Jul 2026

DP-MGTD: Privacy-Preserving Machine-Generated Text Detection via Adaptive Differentially Private Entity Sanitization

ResearchDGX agent

arXiv:2601.04641v2 Announce Type: replace-cross Abstract: The deployment of Machine-Generated Text (MGT) detection systems necessitates processing sensitive user data, creating a fundamental conflict

Dual-Path LLM Reasoning for Multimodal Few-Shot Knowledge Graph Completion

ApplicationsDGX agent

arXiv:2607.26909v1 Announce Type: new Abstract: Knowledge graph completion (KGC) aims to infer missing facts in knowledge graphs (KGs), thereby improving their completeness and supporting downstream i

DuplexGen: Adaptive Synthesis of Human-AI Turn-Taking Dialogues

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2607.26178v1 Announce Type: new Abstract: Turn-taking is a central component of full-duplex interaction. Which turn-taking behaviors are appropriate varies with the scenario, yet current models

Enhancing Generative Information Extraction with Two-step Validation: A Product Attribute Use Case

Model ReleasesDGX agent

arXiv:2607.26780v1 Announce Type: new Abstract: The ability of large language models (LLMs) to process and generate text has introduced potential for applications in information extraction (IE). While

Evaluating Prompt Scope and Demonstration Similarity in Local LLM Machine Translation

Model ReleasesDGX agent

arXiv:2607.26286v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as general-purpose translation systems, but their behavior is usually evaluated under a single prompt

Evaluating Regional Bias in LLMs From Abstract Stereotype to Concrete Social Decision-Making

SafetyDGX agent

arXiv:2607.27022v1 Announce Type: new Abstract: Regional bias in large language models (LLMs) may shape both perceptions of regional groups and decisions about individuals from different regions. Yet

exttt{AMEND++}: Benchmarking Eligibility Criteria Amendments in Clinical Trials

Model ReleasesDGX agent

arXiv:2601.06300v2 Announce Type: replace Abstract: Clinical trial amendments frequently introduce delays, increased costs, and administrative burden, with eligibility criteria being the most commonly

FedWeave: Rethinking the Unit of Specialization in Heterogeneous Federated MoE-LoRA

Model ReleasesDGX agent

arXiv:2607.26618v1 Announce Type: cross Abstract: Federated PEFT enables LLMs to collaboratively adapt to decentralized private data without sharing raw examples. However, task heterogeneity across cl

Filesystem-Based Memory for LLM Agents: Organization, Evolution, and Sustainability

AgentsDGX agent

arXiv:2607.26637v1 Announce Type: new Abstract: Deployed LLM agents increasingly keep their long-term memory as a filesystem: a directory tree of markdown files that the agent itself reads, writes, an

ForgetBench: Benchmarking Forgetting Dynamics of Long-Term Parametric Memory in Language Models

Model ReleasesDGX agent

arXiv:2607.26455v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated strong capabilities in knowledge acquisition and reasoning, yet their ability to retain previously acquir

From Found to Designed: Concepts as a Design Axis for Large Language Models

SafetyDGX agent

arXiv:2607.26825v1 Announce Type: new Abstract: Large language models (LLMs) encode rich concept-like information, but represent it implicitly through distributed statistical associations rather than

From Representations to Behaviors: Exploring the Person-Situation-Behavior Triad in LLMs

ResearchDGX agent

arXiv:2607.26853v1 Announce Type: new Abstract: Human personality theories characterize traits not as isolated attributes captured by a single score, but as stable individual tendencies expressed thro

Generation or Judgement? A Paradigm Perspective on LLM-Based Emotion-Cause Pair Extraction in Conversation

ResearchDGX agent

arXiv:2607.26967v1 Announce Type: new Abstract: Emotion-cause pair extraction in conversation (ECPEC) identifies utterance pairs in which one utterance causes an emotion expressed in another. Recent L

GPT-Red: Automated Red Teaming via Self-Play at Scale

Model ReleasesDGX agent

arXiv:2607.26115v1 Announce Type: cross Abstract: We introduce extbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal

GroupRAG: Cognitively Inspired Group-Aware Retrieval and Reasoning via Knowledge-Driven Problem Structuring

ApplicationsDGX agent

arXiv:2603.26807v2 Announce Type: replace-cross Abstract: The performance of language models is commonly limited by insufficient knowledge and constrained reasoning. Prior approaches such as Retrieval

Hearsay: Vision-Language Medical Diagnoses Without an Image

Model ReleasesDGX agent

arXiv:2607.26886v1 Announce Type: cross Abstract: When asked to describe a medical image that was never attached, frontier vision-language models do not abstain: they confabulate a diagnosis. We show

Heterogeneous Element-Aware Cross-Version Differencing of Scientific Documents via Layout-Aware Alignment and Structure-Aware Reasoning

SafetyDGX agent

arXiv:2607.14117v2 Announce Type: replace Abstract: Cross-version differencing of scientific documents is essential in scholarly publishing and technical documentation, but remains challenging because

Identifying Implicit Bias in LLM-based Chat AI Toward People with Intellectual Disabilities

Model ReleasesDGX agent

arXiv:2607.26062v1 Announce Type: cross Abstract: Background: This work investigates the presence of implicit bias in Large Language Model (LLM)-based chat AI models directed toward people with intell

(Im)Paired Programming: Coding Agents Improve Productivity but Harm Understanding

AgentsDGX agent

arXiv:2607.26375v1 Announce Type: new Abstract: Coding agents (e.g., Cursor) improve developer productivity by optimizing task completion, but shifting users from writing code to prompting and reviewi

Keyless Attention: Value-Space Routing and Value-Only Caching for Efficient Transformers

ResearchDGX agent

arXiv:2606.21848v2 Announce Type: replace Abstract: Transformer architectures form the foundation of modern natural language processing, making it crucial to address the efficiency and scalability lim

Knowledge before Reasoning: EC-Reason-Bench, a Training-Free Diagnostic Benchmark for LLM Enzyme Classification

Model ReleasesDGX agent

arXiv:2607.26397v1 Announce Type: new Abstract: Enzyme function prediction is a hierarchical, knowledge-intensive form of protein function classification. Existing benchmarks expose an anomaly: genera

LAMUS: A Large-Scale Corpus for Legal Argument Mining from U.S. Caselaw using LLMs

ApplicationsDGX agent

arXiv:2603.08286v2 Announce Type: replace Abstract: Legal argument mining aims to identify and classify the functional components of judicial reasoning, such as facts, issues, rules, analysis, and con

Language Models are not Equally Robust to Non-Canonical Tokenization across Languages

Model ReleasesDGX agent

arXiv:2607.26831v1 Announce Type: new Abstract: Despite the existence of exponentially many valid tokenizations for a given string, language models operate on a single canonical sequence deterministic

Large Emotional World Model

ResearchDGX agent

arXiv:2512.24149v2 Announce Type: replace Abstract: The world is governed by both physical laws and affective dynamics. Physical laws govern state transitions, while affective dynamics shape human act

Large-Scale ChatBot Validation Through Customer Digital Twin Simulations

SafetyDGX agent

arXiv:2607.26060v1 Announce Type: new Abstract: LLM-based chatbots are transforming customer service in regulated domains such as banking, but scalable and cost-effective validation remains a critical

Latent-IM: Latent Interaction Management for Speech LLMs

SafetyDGX agent

arXiv:2607.26928v1 Announce Type: new Abstract: Classical spoken dialogue systems often separated dialogue management from response realization: a policy selected the next dialogue action, and a gener

Learning Dynamic User Personas from Implicit Interaction Streams via Iterative Refinement

ApplicationsDGX agent

arXiv:2607.26473v1 Announce Type: cross Abstract: Personalizing large language models (LLMs) to individual users is essential for improving user experience, yet existing approaches typically rely on e

Linguistic Monoculture in LLM-Assisted Language Use

ResearchDGX agent

arXiv:2607.27134v1 Announce Type: cross Abstract: Writing and communication are increasingly mediated by large language models (LLMs) that are being used to draft, revise and polish text. Although suc

Living-Harness Is an Interactive-Agent Evolver

AgentsDGX agent

arXiv:2607.26598v1 Announce Type: cross Abstract: Large language model (LLM) agents may recover from a failure within an episode or after a retry, yet the same execution failure can recur in later tas

Making Implicit Premises Explicit in Logical Understanding of Enthymemes

ApplicationsDGX agent

arXiv:2603.06114v2 Announce Type: replace Abstract: Real-world arguments in text and dialogues are normally enthymemes (i.e. some of their premises and/or claims are implicit). Natural language proces

MediaWiki Code2Code Search: Neural Retrieval for the Semantic Discovery of Open-Source Software Entities

Model ReleasesDGX agent

arXiv:2607.26766v1 Announce Type: cross Abstract: Code search in large-scale ecosystems is often hindered by the lexical gap between user queries and implementation details, alongside the trade-off be

MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications

Model ReleasesDGX agent

arXiv:2409.07314v4 Announce Type: replace Abstract: While Large Language Models (LLMs) achieve superhuman performance on standardized medical licensing exams, these static benchmarks have become satur

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent

AgentsDGX agent

arXiv:2507.02259v2 Announce Type: replace Abstract: Despite improvements by length extrapolation, efficient attention and memory modules, handling infinitely long documents with linear complexity with

Mental World Modeling

AgentsDGX agent

arXiv:2607.27201v1 Announce Type: new Abstract: World models enable a predictive substrate for planning and action, yet existing formulations merely answer a physical question: what/where it is, and h

Mergeable Model-Side Aggregation States for Long-Context Language Models

Model ReleasesDGX agent

arXiv:2607.26448v1 Announce Type: new Abstract: A known limitation of long-context language models is their increasingly unreliable performance in non-additive, set-based aggregation as context length

Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback

Model ReleasesDGX agent

arXiv:2607.26094v1 Announce Type: cross Abstract: Reinforcement Learning from Human Feedback (RLHF) is the standard approach for aligning large language models with human preferences, but its quality

Metis: Memory Foundation Model

AgentsDGX agent

arXiv:2607.26760v1 Announce Type: new Abstract: Recent advances in AI agents have increasingly internalized native capabilities into their underlying foundation models, giving rise to multimodal found

MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis

AgentsDGX agent

arXiv:2607.27146v1 Announce Type: cross Abstract: Coding agents have made substantial progress on software engineering tasks that modify existing codebases, including bug fixing and feature implementa

Misalignment Has a Personality: A Big Five Account of Emergent Misalignment

SafetyDGX agent

arXiv:2607.26389v1 Announce Type: new Abstract: Fine-tuning a language model on data containing a narrow flaw, such as insecure code or incorrect mathematical answers, can cause broad misalignment thr

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation

Model ReleasesDGX agent

arXiv:2509.22768v3 Announce Type: replace Abstract: We introduce ML2B, the first benchmark for evaluating cross-lingual task comprehension in end-to-end ML pipeline generation by large language models

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

Model ReleasesDGX agent

arXiv:2607.27155v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support

On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment

SafetyDGX agent

arXiv:2607.27081v1 Announce Type: cross Abstract: Fine-tuning is the dominant paradigm for specializing large language models (LLMs), yet it exposes a critical vulnerability: malicious data providers

OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment

SafetyDGX agent

arXiv:2607.26981v1 Announce Type: new Abstract: Large language models are increasingly used as decision aids whose probability judgments shape downstream choices. Whether those judgments carry a syste

Pangram 4 Technical Report

ResearchDGX agent

arXiv:2607.27183v1 Announce Type: new Abstract: We present Pangram 4, the latest deep-learning-based AI-text classification model from Pangram Labs. We achieve an AUROC of 0.9916 with a false positive

Phoneme- vs. Character-Level Targets and Selective State-Space Models for Intracortical Brain-to-Text

Model ReleasesDGX agent

arXiv:2607.26751v1 Announce Type: new Abstract: State-of-the-art intracortical brain-to-text systems pair a neural-sequence phone decoder with an external language model. Two design axes remain undere

Position: Evaluation Scores Are Perishable Knowledge Claims

Model ReleasesDGX agent

arXiv:2607.26191v1 Announce Type: cross Abstract: Evaluation methodologies for language models increasingly combine multiple signals, from automated metrics and LLM-as-judge ratings to human assessmen

Probing the Origins of Reasoning Performance: Representational Quality for Mathematical Problem-Solving in RL vs. SFT Fine-Tuned Models

SafetyDGX agent

arXiv:2607.26119v1 Announce Type: cross Abstract: Large reasoning models trained via reinforcement learning (RL) have been increasingly shown to outperform their supervised fine-tuned (SFT) counterpar

Prosody-driven Jailbreaks in Audio LLMs: A Controlled Study and Mechanistic Analysis

Model ReleasesDGX agent

arXiv:2607.26541v1 Announce Type: cross Abstract: Audio-capable foundation models enable end-to-end spoken interaction, but they also introduce safety risks beyond transcript content. It remains uncle

Relation Geometry in Semantic Space of Language Models

ResearchDGX agent

arXiv:2607.26762v1 Announce Type: new Abstract: When it comes to generating vector representations of words, current language models are achieving high-quality results. However, what is not known is t

Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes

ResearchDGX agent

arXiv:2607.26627v1 Announce Type: new Abstract: Speculative Decoding (SD) accelerates large language model inference by allowing a lightweight draft model to propose tokens that are subsequently verif

Robostreet Flow: A Lightweight, Ultra-Low-Drag Electric Tractor and Four-Truck Hybrid Convoy Architecture for Minimum-Cost Point-to-Point Freight

SafetyDGX agent

arXiv:2607.26250v1 Announce Type: new Abstract: Line-haul trucking costs are dominated by three comparably sized components: energy, driver labor, and equipment. Most efficiency technologies address o

Same Evidence, Different Target: Decoding How Diagnostic Evidence Bears on Causal Questions from Language-Model States

Model ReleasesDGX agent

arXiv:2607.26929v1 Announce Type: new Abstract: The same diagnostic result can support or challenge one causal claim yet fail to address another when the claims concern different populations, outcomes

Scientific Knowledge Discovery in the Age of Large Language Models

ResearchDGX agent

arXiv:2607.26670v1 Announce Type: cross Abstract: The rapid growth of scholarly literature has made identifying relevant publications increasingly difficult, and conventional search systems still depe

SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response

Model ReleasesDGX agent

arXiv:2607.26791v1 Announce Type: cross Abstract: Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning

Model ReleasesDGX agent

arXiv:2607.26873v1 Announce Type: new Abstract: Test-time reinforcement learning (TTRL) enables language models to self-evolve at inference time without labeled feedback. Existing methods rely on answ

Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data

Model ReleasesDGX agent

arXiv:2607.27056v1 Announce Type: cross Abstract: Personalized agents are increasingly applied to assist users across a wide range of tasks. Effective personalized assistance requires not only retriev

SkillCAT: Contrastive, Assessment-Augmented and Topology-AwareSkill Self-Evolution for LLM Agents

AgentsDGX agent

arXiv:2606.13317v2 Announce Type: replace Abstract: Skill self-evolution methods for LLM agents aim to turn execution trajectories into reusable skill documents. However, current pipelines typically d

SpecFirst: Behavioral Specification Elicitation as a First-Class Step in Agent-Based Program Synthesis from Scratch

AgentsDGX agent

arXiv:2607.27167v1 Announce Type: cross Abstract: LLM-based agents excel at software engineering tasks where an existing codebase provides context, but constructing a program from scratch remains fund

Statistical laws and linguistics differ in naturalistic video and fictional conversations

ResearchDGX agent

arXiv:2512.18072v3 Announce Type: replace Abstract: Conversation is a cornerstone of social connection and is linked to well-being outcomes. Conversations vary widely in type with some portion generat

Steering Instruction Hierarchies at Inference Time

SafetyDGX agent

arXiv:2607.26228v1 Announce Type: new Abstract: Instruction hierarchies are a core safety assumption of language model deployment: higher priority inputs, such as system prompts, should override confl

← Previous
1…1415161718…128
Next →