AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
23 Jul 2026

On the Computational Complexity of Structural Generalization

Model ReleasesDGX agent

arXiv:2607.19573v1 Announce Type: new Abstract: Structural generalization has been measured repeatedly by several benchmarks, yet it has never been formally defined. We give a definition that translat

OpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party Skills

Model ReleasesDGX agent

arXiv:2607.20121v1 Announce Type: new Abstract: LLM-based agents leverage third-party skills to extend their capabilities in open-world scenarios. However, third-party skills can introduce extra secur

PyroDash: Cost-Efficient Token-Level Small-Large Language Model Collaborative Inference

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2607.20327v1 Announce Type: new Abstract: Large language models (LLMs) provide strong reasoning capabilities but are expensive to serve at scale, whereas small language models (SLMs) are cheaper

RALS: Resources and Baselines for Romanian Automatic Lexical Simplification

ResearchDGX agent

arXiv:2607.20078v1 Announce Type: new Abstract: We introduce the first dataset that jointly covers both lexical complexity prediction (LCP) annotations and lexical simplification (LS) for Romanian, al

Scaling Laws for Hypernetwork-Based Knowledge Injection in Large Language Models

TutorialsDGX agent

arXiv:2607.19604v1 Announce Type: new Abstract: Injecting factual knowledge into large language models (LLMs) reliably and at scale remains an open challenge. Hypernetworks provide a promising solutio

Simultaneous Speech-to-Speech Translation Without Aligned Data

Model ReleasesDGX agent

arXiv:2602.11072v2 Announce Type: replace Abstract: Simultaneous speech translation requires translating source speech into a target language in real-time while handling non-monotonic word dependencie

Solar Open 2 Technical Report

Model ReleasesDGX agent

arXiv:2607.20062v1 Announce Type: new Abstract: We present Solar Open 2, a 250B-A15B Mixture-of-Experts language model built for long-horizon agentic tasks, scaled up from Solar Open 1 (Solar Open 100

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning

SafetyDGX agent

arXiv:2607.18722v2 Announce Type: replace-cross Abstract: Asynchronous reinforcement learning improves throughput by decoupling rollout generation from optimization, but staleness is an inevitable byp

STEMTOX: From Collaborative Tags to Fine-Grained Toxic Meme Detection via Entropy-Guided Multi-Task Learning

ApplicationsDGX agent

arXiv:2508.04166v3 Announce Type: replace-cross Abstract: Memes, as a widely used mode of online communication, often serve as vehicles for spreading harmful content. However, limitations in data acce

surprisal is Not a Theory

ResearchDGX agent

arXiv:2607.20208v1 Announce Type: new Abstract: Surprisal Theory is often characterized as a computational-level explanation per (Marr, 1982). We argue in this work that, even though a computational l

TalentCLEF at CLEF2026: Skill and Job Title Intelligence for Human Capital Management

SafetyDGX agent

arXiv:2607.20009v1 Announce Type: new Abstract: This paper presents the second edition of the TalentCLEF Challenge, which will run as an evaluation lab as part of CLEF 2026. The aim of TalentCLEF is t

Task Competence Is Not Instruction Following: Evaluating Instruction-Conflicting Behavior in Small Language Models

Model ReleasesDGX agent

arXiv:2607.19608v1 Announce Type: new Abstract: Instruction tuning is meant to make language models follow user requests, yet it is unclear whether small models comply when an instruction conflicts wi

Test-Time Training for Modality Order Consistency in Vision-Language Models

ResearchDGX agent

arXiv:2607.20351v1 Announce Type: cross Abstract: We find that vision-language models are sensitive to a specific semantically irrelevant change: the order in which the image and question are presente

The Blessing of Dimensionality: How Near-Orthogonality in High-Dimensional Spaces Explains Temporal Portability

Model ReleasesDGX agent

arXiv:2607.20301v1 Announce Type: cross Abstract: Fine-tuning has been widely used to adapt large language models (LLMs) for domain-specific tasks. Parameter efficient fine-tuning (PEFT) methods such

The Two-Process Theory of Machine Self-Report

SafetyDGX agent

arXiv:2607.20082v1 Announce Type: new Abstract: Language models are increasingly asked to self-report, informing safety evaluations, public understanding, and model-welfare debates. Yet their reports

TriAgent: Divergence-Aware Multi-Agent Committees for Cost-Efficient Financial Sentiment Analysis

Model ReleasesDGX agent

arXiv:2607.19794v1 Announce Type: new Abstract: Production LLM-based financial sentiment analysis faces a structural cost trap: most queries are trivially classifiable, yet expensive cloud reasoners p

Twin Agent: Context Residual Compression for Privilege Separated Agents

AgentsDGX agent

arXiv:2607.19595v1 Announce Type: cross Abstract: Large language model (LLM) agents are vulnerable to security risks, such as prompt injection attacks from untrusted context that manipulate downstream

Two-Step Occupation Coding

ResearchDGX agent

arXiv:2607.20101v1 Announce Type: new Abstract: Occupation coding links job titles in free text to occupational taxonomies and is a core task in labor market research. Existing approaches typically ad

Understanding the Impact of Linguistic Realization Choices on LLM Stance with Causal Tracing

Local AiDGX agent

arXiv:2607.20115v1 Announce Type: new Abstract: Large language models (LLMs) are known to be sensitive to prompt and input formulations. However, existing studies have focused on lexical realization a

VizRAG: Enhancing Retrieval-Augmented Generation with Hypergraph Visualization

ResearchDGX agent

arXiv:2607.19830v1 Announce Type: new Abstract: Hypergraph-based RAG systems surpass traditional graph-based approaches by organizing complex n-ary atomic facts among entities, rather than relying sol

When Benchmarks Mislead: Shortcut Learning, Length Confounds, and the Limits of Cross-Dataset Generalization in Multilingual Fake News and Sarcasm Detection

ApplicationsDGX agent

arXiv:2607.14131v2 Announce Type: replace Abstract: Cross-dataset generalisation is a fundamental requirement for deploying text classifiers in real-world settings, yet systematic evaluation across co

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play

SafetyDGX agent

arXiv:2607.19523v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) is widely used to adapt large language models to downstream tasks, but its effect on behavioral diversity in sequential dec

Which Values Do LLMs Confuse? A Schwartz-Based Recognition Study

SafetyDGX agent

arXiv:2607.20270v1 Announce Type: new Abstract: Large language models are increasingly evaluated through the values they endorse, but such evaluations presuppose that models can identify the value exp

15 Jul 2026

A JoLT for the KV Cache: Near-Lossless KV Cache Compression via Joint Tucker and JL-Residual Allocation for LLMs

Model ReleasesDGX agent

arXiv:2607.12550v1 Announce Type: cross Abstract: The key-value (KV) cache has become the dominant memory cost of transformer inference. It grows with batch size, context length, and depth, and at lon

A Shared Subcircuit Lets LLMs Count Down Across Tasks

Model ReleasesDGX agent

arXiv:2607.12279v1 Announce Type: new Abstract: Writing a sentence of exactly twelve words; ending a DNA sequence at the right codon; formatting an ASCII table. These are all tasks that language model

Agentic systems for breast cancer treatment recommendations

Model ReleasesDGX agent

arXiv:2607.12051v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly being explored for clinical decision support, but their reliability in complex oncology treatment planning

Amplitude-Only FFN Intervention for Tool-Structured LLM Inference Method: Gated Evaluation Protocol, and Cross-Model Empirical Results

AgentsDGX agent

arXiv:2607.11183v2 Announce Type: replace Abstract: Large language models increasingly operate as tool-using agents, where small format, argument, or function-call errors can invalidate otherwise plau

Belief-reality separation lives in routing over a shared value slot in language models

Model ReleasesDGX agent

arXiv:2607.11945v1 Announce Type: new Abstract: Capable language models hold what a character believes apart from what is true: told 'Anna believes the cup is blue; in reality it is red,' they answer

Beyond Binary Detection: A Multi-Dimensional Taxonomy of Cancer Misinformation on Reddit

ResearchDGX agent

arXiv:2607.12383v1 Announce Type: new Abstract: Cancer-related discussions on social media provide an important space for information exchange and peer support, but also facilitate the spread of misin

Beyond Parallel Tracking: Interactive Multi-Feature Fusion Drives Semantic Reconstruction from Non-invasive Brain Recordings

SafetyDGX agent

arXiv:2607.12071v1 Announce Type: new Abstract: Continuous semantic reconstruction from non-invasive neural recordings remains limited by the representational mismatch between semantic feature spaces

Can a Language Model Learn Facts Continually in Its Weights?

TutorialsDGX agent

arXiv:2607.11020v2 Announce Type: replace Abstract: Continual learning promises a language model that keeps acquiring knowledge after training, with each new fact written into its weights. Whether wei

Can LLMs Write Reliable Rubrics? A Meta-Evaluation for Experiment Reproduction

SafetyDGX agent

arXiv:2607.12835v1 Announce Type: new Abstract: Rubric-based evaluation is a promising approach for assessing open-ended outputs from LLM-based research agents, particularly in paper reproduction, whe

CityBehavEx: A Scalable and Empirically Validated LLM-Assisted Urban Simulation Platform

SafetyDGX agent

arXiv:2607.12086v1 Announce Type: new Abstract: Recent LLM-based multi-agent urban simulators can generate semantically rich city routines, but they remain costly to scale and are often weakly validat

Entropy in Semantic Memory Navigation in Blind and Sighted Individuals: The Effect of Visual Experience

ResearchDGX agent

arXiv:2607.12185v1 Announce Type: new Abstract: Embodied accounts of semantic memory highlight the role of sensorimotor systems in acquiring and storing knowledge. Congenitally blind populations offer

Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning Models

ResearchDGX agent

arXiv:2602.02244v3 Announce Type: replace-cross Abstract: The standard post-training recipe for large reasoning models, supervised fine-tuning followed by reinforcement learning (SFT-then-RL), may lim

Epistemic Stance Flexibility Probing: Measuring Prompt-Conditioned Register Shift in Large Language Models

Model ReleasesDGX agent

arXiv:2607.12739v1 Announce Type: new Abstract: A language model may be asked either what experts believe about a contested claim or what it believes about the claim itself. A trustworthy conversation

Evaluating Large Language Models on Misconceptions in Multi-Turn Medical Conversations

Model ReleasesDGX agent

arXiv:2607.12884v1 Announce Type: new Abstract: Patients seeking medical information often ask questions that embed incorrect assumptions or misconceptions. In such cases, safe medical communication r

Extractable Memorization From First Principles

Model ReleasesDGX agent

arXiv:2607.12649v1 Announce Type: cross Abstract: Recent work on extractable memorization in LLMs suffers from two contrasting validity problems. Some studies overstate extraction, e.g., relying on se

FairCoder: Probing LLM Bias in High-Stakes Decision Making via Coding Tasks

Model ReleasesDGX agent

arXiv:2501.05396v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used in high-stakes decisions such as hiring and college admissions, making their social bias a critic

Fine-Tuned Multi-Agent Framework for Detecting OCEAN in Life Narratives

AgentsDGX agent

arXiv:2607.12215v1 Announce Type: new Abstract: Accurately assessing personality from text is challenging because traits are latent, context-dependent, and often subtly expressed across long narrative

FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality

Model ReleasesDGX agent

arXiv:2607.12252v1 Announce Type: new Abstract: Deep research agents are increasingly used to produce long-form financial reports, yet large-scale evaluation remains bottlenecked by the need for human

From Sentiment to Actionable Insights: Public Sentiment Analysis of Advanced Air Mobility

SafetyDGX agent

arXiv:2606.20751v2 Announce Type: replace Abstract: Advanced Air Mobility (AAM) is an emerging low-altitude transportation system whose successful deployment depends on both technological progress and

From Words to Widgets for Controllable LLM Generation

ResearchDGX agent

arXiv:2604.10925v2 Announce Type: cross Abstract: Natural language remains the predominant way people interact with large language models (LLMs). However, users often struggle to precisely express and

Growing a Tail: Increasing Output Diversity in Large Language Models

SafetyDGX agent

arXiv:2411.02989v2 Announce Type: replace Abstract: How diverse are the outputs of large language models when diversity is desired? We examine the diversity of responses of several language models to

Hierarchical Latent Structures in Data Generation Process Unify Mechanistic Phenomena across Scale

ResearchDGX agent

arXiv:2603.06592v2 Announce Type: replace Abstract: Contemporary studies in mechanistic interpretability have uncovered many puzzling phenomena in the neural information processing of Transformer-base

Hybrid Continual Learning for Low-Resource Australian Aboriginal Language Identification

ResearchDGX agent

arXiv:2607.11946v1 Announce Type: new Abstract: Language identification is an important step toward integrating endangered Australian Aboriginal languages (AALs) into speech technologies supporting la

KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill

Model ReleasesDGX agent

arXiv:2607.12625v1 Announce Type: new Abstract: OpenClaw has emerged as a leading agent framework for complex task automation, yet it faces insufficient cross-platform GUI interaction support and a we

Knowledgeless Language Models: Suppressing Parametric Recall for Evidence-Grounded Language Modeling

ResearchDGX agent

arXiv:2607.12831v1 Announce Type: new Abstract: Language models encode substantial factual knowledge in their parameters, which can lead to unreliable behavior when this knowledge is outdated, incompl

Language Identification with Succinct Machine-Independent Traces

TutorialsDGX agent

arXiv:2607.12443v1 Announce Type: new Abstract: Motivated by the power of large language models, there has been renewed interest in the Gold-Angluin model of language identification in the limit, with

Learning Mechanistic Reasoning for Chemical Reactions with Large Language Models

Model ReleasesDGX agent

arXiv:2607.12771v1 Announce Type: cross Abstract: Reaction mechanisms consist of the step-by-step sequences of elementary reactions that explain chemical transformations. Learning the mechanism logic

LLM Judges Can Be Too Generous When There Is No Reference Answer

ResearchDGX agent

arXiv:2607.12885v1 Announce Type: new Abstract: LLM judges are increasingly being used to evaluate open-ended model responses, often in no-reference settings where a ground-truth answer is unavailable

Lost in the Maze: Overcoming Context Limitations in Long-Horizon Agentic Search

Model ReleasesDGX agent

arXiv:2510.18939v2 Announce Type: replace Abstract: Long-horizon agentic search requires iteratively exploring the web over long trajectories and synthesizing information across many sources, enabling

MAGE: Understanding Stability-Performance Trade-offs in Multi-component Prompt Optimization

Model ReleasesDGX agent

arXiv:2607.11944v1 Announce Type: new Abstract: How do different components of iterative prompt optimization interact, and what happens when they are combined? We investigate this through MAGE (Memory

Policy-Conditioned Constrained Decoding for Column-Level Access Control in Text-to-SQL

SafetyDGX agent

arXiv:2607.12341v1 Announce Type: new Abstract: Text-to-SQL is increasingly deployed across trust boundaries between data providers and users. Such deployment must balance three competing requirements

Predict the Retrieval! Test time adaptation for Retrieval Augmented Generation

Model ReleasesDGX agent

arXiv:2601.11443v3 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) has emerged as a powerful approach for enhancing large language models' question-answering capabilities through

QUBO-Optimized Evidence Selection for Retrieval-Augmented Question Answering with Unconventional Solvers

ResearchDGX agent

arXiv:2607.12334v1 Announce Type: new Abstract: Retrieval-augmented question answering depends on selecting evidence passages that jointly support answer generation. However, many RAG pipelines rely o

Rethinking Evaluation in Retrieval-Augmented Personalized Dialogue: A Cognitive and Linguistic Perspective

ApplicationsDGX agent

arXiv:2603.14217v3 Announce Type: replace Abstract: In cognitive science and linguistic theory, dialogue is not seen as a chain of independent utterances but rather as a joint activity sustained by co

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning

ResearchDGX agent

arXiv:2607.12395v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards without human-annotated data, often referred to as zero RL, has emerged as a powerful paradigm for elicit

Segregate, Refine, Integrate: Decomposing Multimodal Fusion for Sentiment Analysis

ResearchDGX agent

arXiv:2607.12686v1 Announce Type: new Abstract: Multimodal fusion must simultaneously refine modality-specific signals and model cross-modal interactions; two competing objectives typically entangled

Speculate with Memory: Lossless Acceleration for LLM Agents

AgentsDGX agent

arXiv:2607.12236v1 Announce Type: cross Abstract: Speculative execution accelerates LLM agents by using a smaller, cheaper model to predict and pre-launch the next step while the environment is idle.

← Previous
1…2122232425…129
Next →