AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlog
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
Research

Generation or Judgement? A Paradigm Perspective on LLM-Based Emotion-Cause Pair Extraction in Conversation

DGX agent

arXiv:2607.26967v1 Announce Type: new Abstract: Emotion-cause pair extraction in conversation (ECPEC) identifies utterance pairs in which one utterance causes an emotion expressed in another. Recent L

researcharxiv-cs-cl
30 Jul 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

GPT-Red: Automated Red Teaming via Self-Play at Scale

DGX agent

arXiv:2607.26115v1 Announce Type: cross Abstract: We introduce extbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal

model-releasesarxiv-cs-cl
30 Jul 2026
Applications

GroupRAG: Cognitively Inspired Group-Aware Retrieval and Reasoning via Knowledge-Driven Problem Structuring

DGX agent

arXiv:2603.26807v2 Announce Type: replace-cross Abstract: The performance of language models is commonly limited by insufficient knowledge and constrained reasoning. Prior approaches such as Retrieval

applicationsarxiv-cs-cl
30 Jul 2026
Model Releases

Hearsay: Vision-Language Medical Diagnoses Without an Image

DGX agent

arXiv:2607.26886v1 Announce Type: cross Abstract: When asked to describe a medical image that was never attached, frontier vision-language models do not abstain: they confabulate a diagnosis. We show

model-releasesarxiv-cs-cl
30 Jul 2026
Safety

Heterogeneous Element-Aware Cross-Version Differencing of Scientific Documents via Layout-Aware Alignment and Structure-Aware Reasoning

DGX agent

arXiv:2607.14117v2 Announce Type: replace Abstract: Cross-version differencing of scientific documents is essential in scholarly publishing and technical documentation, but remains challenging because

safetyarxiv-cs-cl
30 Jul 2026
Model Releases

Identifying Implicit Bias in LLM-based Chat AI Toward People with Intellectual Disabilities

DGX agent

arXiv:2607.26062v1 Announce Type: cross Abstract: Background: This work investigates the presence of implicit bias in Large Language Model (LLM)-based chat AI models directed toward people with intell

model-releasesarxiv-cs-cl
30 Jul 2026
Agents

(Im)Paired Programming: Coding Agents Improve Productivity but Harm Understanding

DGX agent

arXiv:2607.26375v1 Announce Type: new Abstract: Coding agents (e.g., Cursor) improve developer productivity by optimizing task completion, but shifting users from writing code to prompting and reviewi

agentsarxiv-cs-cl
30 Jul 2026
Research

Keyless Attention: Value-Space Routing and Value-Only Caching for Efficient Transformers

DGX agent

arXiv:2606.21848v2 Announce Type: replace Abstract: Transformer architectures form the foundation of modern natural language processing, making it crucial to address the efficiency and scalability lim

researcharxiv-cs-cl
30 Jul 2026
Model Releases

Knowledge before Reasoning: EC-Reason-Bench, a Training-Free Diagnostic Benchmark for LLM Enzyme Classification

DGX agent

arXiv:2607.26397v1 Announce Type: new Abstract: Enzyme function prediction is a hierarchical, knowledge-intensive form of protein function classification. Existing benchmarks expose an anomaly: genera

model-releasesarxiv-cs-cl
30 Jul 2026
Applications

LAMUS: A Large-Scale Corpus for Legal Argument Mining from U.S. Caselaw using LLMs

DGX agent

arXiv:2603.08286v2 Announce Type: replace Abstract: Legal argument mining aims to identify and classify the functional components of judicial reasoning, such as facts, issues, rules, analysis, and con

applicationsarxiv-cs-cl
30 Jul 2026
Model Releases

Language Models are not Equally Robust to Non-Canonical Tokenization across Languages

DGX agent

arXiv:2607.26831v1 Announce Type: new Abstract: Despite the existence of exponentially many valid tokenizations for a given string, language models operate on a single canonical sequence deterministic

model-releasesarxiv-cs-cl
30 Jul 2026
Research

Large Emotional World Model

DGX agent

arXiv:2512.24149v2 Announce Type: replace Abstract: The world is governed by both physical laws and affective dynamics. Physical laws govern state transitions, while affective dynamics shape human act

researcharxiv-cs-cl
30 Jul 2026
Safety

Large-Scale ChatBot Validation Through Customer Digital Twin Simulations

DGX agent

arXiv:2607.26060v1 Announce Type: new Abstract: LLM-based chatbots are transforming customer service in regulated domains such as banking, but scalable and cost-effective validation remains a critical

safetyarxiv-cs-cl
30 Jul 2026
Safety

Latent-IM: Latent Interaction Management for Speech LLMs

DGX agent

arXiv:2607.26928v1 Announce Type: new Abstract: Classical spoken dialogue systems often separated dialogue management from response realization: a policy selected the next dialogue action, and a gener

safetyarxiv-cs-cl
30 Jul 2026
Applications

Learning Dynamic User Personas from Implicit Interaction Streams via Iterative Refinement

DGX agent

arXiv:2607.26473v1 Announce Type: cross Abstract: Personalizing large language models (LLMs) to individual users is essential for improving user experience, yet existing approaches typically rely on e

applicationsarxiv-cs-cl
30 Jul 2026
Research

Linguistic Monoculture in LLM-Assisted Language Use

DGX agent

arXiv:2607.27134v1 Announce Type: cross Abstract: Writing and communication are increasingly mediated by large language models (LLMs) that are being used to draft, revise and polish text. Although suc

researcharxiv-cs-cl
30 Jul 2026
Agents

Living-Harness Is an Interactive-Agent Evolver

DGX agent

arXiv:2607.26598v1 Announce Type: cross Abstract: Large language model (LLM) agents may recover from a failure within an episode or after a retry, yet the same execution failure can recur in later tas

agentsarxiv-cs-cl
30 Jul 2026
Applications

Making Implicit Premises Explicit in Logical Understanding of Enthymemes

DGX agent

arXiv:2603.06114v2 Announce Type: replace Abstract: Real-world arguments in text and dialogues are normally enthymemes (i.e. some of their premises and/or claims are implicit). Natural language proces

applicationsarxiv-cs-cl
30 Jul 2026
Model Releases

MediaWiki Code2Code Search: Neural Retrieval for the Semantic Discovery of Open-Source Software Entities

DGX agent

arXiv:2607.26766v1 Announce Type: cross Abstract: Code search in large-scale ecosystems is often hindered by the lexical gap between user queries and implementation details, alongside the trade-off be

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications

DGX agent

arXiv:2409.07314v4 Announce Type: replace Abstract: While Large Language Models (LLMs) achieve superhuman performance on standardized medical licensing exams, these static benchmarks have become satur

model-releasesarxiv-cs-cl
30 Jul 2026
Agents

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent

DGX agent

arXiv:2507.02259v2 Announce Type: replace Abstract: Despite improvements by length extrapolation, efficient attention and memory modules, handling infinitely long documents with linear complexity with

agentsarxiv-cs-cl
30 Jul 2026
Agents

Mental World Modeling

DGX agent

arXiv:2607.27201v1 Announce Type: new Abstract: World models enable a predictive substrate for planning and action, yet existing formulations merely answer a physical question: what/where it is, and h

agentsarxiv-cs-cl
30 Jul 2026
Model Releases

Mergeable Model-Side Aggregation States for Long-Context Language Models

DGX agent

arXiv:2607.26448v1 Announce Type: new Abstract: A known limitation of long-context language models is their increasingly unreliable performance in non-additive, set-based aggregation as context length

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback

DGX agent

arXiv:2607.26094v1 Announce Type: cross Abstract: Reinforcement Learning from Human Feedback (RLHF) is the standard approach for aligning large language models with human preferences, but its quality

model-releasesarxiv-cs-cl
30 Jul 2026
Agents

Metis: Memory Foundation Model

DGX agent

arXiv:2607.26760v1 Announce Type: new Abstract: Recent advances in AI agents have increasingly internalized native capabilities into their underlying foundation models, giving rise to multimodal found

agentsarxiv-cs-cl
30 Jul 2026
Agents

MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis

DGX agent

arXiv:2607.27146v1 Announce Type: cross Abstract: Coding agents have made substantial progress on software engineering tasks that modify existing codebases, including bug fixing and feature implementa

agentsarxiv-cs-cl
30 Jul 2026
Safety

Misalignment Has a Personality: A Big Five Account of Emergent Misalignment

DGX agent

arXiv:2607.26389v1 Announce Type: new Abstract: Fine-tuning a language model on data containing a narrow flaw, such as insecure code or incorrect mathematical answers, can cause broad misalignment thr

safetyarxiv-cs-cl
30 Jul 2026
Model Releases

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation

DGX agent

arXiv:2509.22768v3 Announce Type: replace Abstract: We introduce ML2B, the first benchmark for evaluating cross-lingual task comprehension in end-to-end ML pipeline generation by large language models

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

DGX agent

arXiv:2607.27155v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support

model-releasesarxiv-cs-cl
30 Jul 2026
Safety

On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment

DGX agent

arXiv:2607.27081v1 Announce Type: cross Abstract: Fine-tuning is the dominant paradigm for specializing large language models (LLMs), yet it exposes a critical vulnerability: malicious data providers

safetyarxiv-cs-cl
30 Jul 2026
Safety

OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment

DGX agent

arXiv:2607.26981v1 Announce Type: new Abstract: Large language models are increasingly used as decision aids whose probability judgments shape downstream choices. Whether those judgments carry a syste

safetyarxiv-cs-cl
30 Jul 2026
Research

Pangram 4 Technical Report

DGX agent

arXiv:2607.27183v1 Announce Type: new Abstract: We present Pangram 4, the latest deep-learning-based AI-text classification model from Pangram Labs. We achieve an AUROC of 0.9916 with a false positive

researcharxiv-cs-cl
30 Jul 2026
Model Releases

Phoneme- vs. Character-Level Targets and Selective State-Space Models for Intracortical Brain-to-Text

DGX agent

arXiv:2607.26751v1 Announce Type: new Abstract: State-of-the-art intracortical brain-to-text systems pair a neural-sequence phone decoder with an external language model. Two design axes remain undere

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

Position: Evaluation Scores Are Perishable Knowledge Claims

DGX agent

arXiv:2607.26191v1 Announce Type: cross Abstract: Evaluation methodologies for language models increasingly combine multiple signals, from automated metrics and LLM-as-judge ratings to human assessmen

model-releasesarxiv-cs-cl
30 Jul 2026
Safety

Probing the Origins of Reasoning Performance: Representational Quality for Mathematical Problem-Solving in RL vs. SFT Fine-Tuned Models

DGX agent

arXiv:2607.26119v1 Announce Type: cross Abstract: Large reasoning models trained via reinforcement learning (RL) have been increasingly shown to outperform their supervised fine-tuned (SFT) counterpar

safetyarxiv-cs-cl
30 Jul 2026
Model Releases

Prosody-driven Jailbreaks in Audio LLMs: A Controlled Study and Mechanistic Analysis

DGX agent

arXiv:2607.26541v1 Announce Type: cross Abstract: Audio-capable foundation models enable end-to-end spoken interaction, but they also introduce safety risks beyond transcript content. It remains uncle

model-releasesarxiv-cs-cl
30 Jul 2026
Research

Relation Geometry in Semantic Space of Language Models

DGX agent

arXiv:2607.26762v1 Announce Type: new Abstract: When it comes to generating vector representations of words, current language models are achieving high-quality results. However, what is not known is t

researcharxiv-cs-cl
30 Jul 2026
Research

Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes

DGX agent

arXiv:2607.26627v1 Announce Type: new Abstract: Speculative Decoding (SD) accelerates large language model inference by allowing a lightweight draft model to propose tokens that are subsequently verif

researcharxiv-cs-cl
30 Jul 2026
Safety

Robostreet Flow: A Lightweight, Ultra-Low-Drag Electric Tractor and Four-Truck Hybrid Convoy Architecture for Minimum-Cost Point-to-Point Freight

DGX agent

arXiv:2607.26250v1 Announce Type: new Abstract: Line-haul trucking costs are dominated by three comparably sized components: energy, driver labor, and equipment. Most efficiency technologies address o

safetyarxiv-cs-cl
30 Jul 2026
Model Releases

Same Evidence, Different Target: Decoding How Diagnostic Evidence Bears on Causal Questions from Language-Model States

DGX agent

arXiv:2607.26929v1 Announce Type: new Abstract: The same diagnostic result can support or challenge one causal claim yet fail to address another when the claims concern different populations, outcomes

model-releasesarxiv-cs-cl
30 Jul 2026
Research

Scientific Knowledge Discovery in the Age of Large Language Models

DGX agent

arXiv:2607.26670v1 Announce Type: cross Abstract: The rapid growth of scholarly literature has made identifying relevant publications increasingly difficult, and conventional search systems still depe

researcharxiv-cs-cl
30 Jul 2026
Model Releases

SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response

DGX agent

arXiv:2607.26791v1 Announce Type: cross Abstract: Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning

DGX agent

arXiv:2607.26873v1 Announce Type: new Abstract: Test-time reinforcement learning (TTRL) enables language models to self-evolve at inference time without labeled feedback. Existing methods rely on answ

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data

DGX agent

arXiv:2607.27056v1 Announce Type: cross Abstract: Personalized agents are increasingly applied to assist users across a wide range of tasks. Effective personalized assistance requires not only retriev

model-releasesarxiv-cs-cl
30 Jul 2026
Agents

SkillCAT: Contrastive, Assessment-Augmented and Topology-AwareSkill Self-Evolution for LLM Agents

DGX agent

arXiv:2606.13317v2 Announce Type: replace Abstract: Skill self-evolution methods for LLM agents aim to turn execution trajectories into reusable skill documents. However, current pipelines typically d

agentsarxiv-cs-cl
30 Jul 2026
Agents

SpecFirst: Behavioral Specification Elicitation as a First-Class Step in Agent-Based Program Synthesis from Scratch

DGX agent

arXiv:2607.27167v1 Announce Type: cross Abstract: LLM-based agents excel at software engineering tasks where an existing codebase provides context, but constructing a program from scratch remains fund

agentsarxiv-cs-cl
30 Jul 2026
Research

Statistical laws and linguistics differ in naturalistic video and fictional conversations

DGX agent

arXiv:2512.18072v3 Announce Type: replace Abstract: Conversation is a cornerstone of social connection and is linked to well-being outcomes. Conversations vary widely in type with some portion generat

researcharxiv-cs-cl
30 Jul 2026
Safety

Steering Instruction Hierarchies at Inference Time

DGX agent

arXiv:2607.26228v1 Announce Type: new Abstract: Instruction hierarchies are a core safety assumption of language model deployment: higher priority inputs, such as system prompts, should override confl

safetyarxiv-cs-cl
30 Jul 2026
← Previous
1…1819202122…160
Next →