AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
Model Releases

Language Diversity: Evaluating Language Usage and AI Performance on African Languages in Digital Spaces

DGX agent

arXiv:2512.01557v3 Announce Type: replace Abstract: This study examines the digital representation of African languages and the challenges this presents for current language detection tools. We evalua

model-releasesarxiv-cs-cl
31 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

Latent States in Neural Networks: Recovering the Temporal Structure of Drifting Data from Model Weights

DGX agent

arXiv:2607.27482v1 Announce Type: cross Abstract: A temporally drifting data stream may pass through discrete regimes rather than changing continuously. We ask whether such regimes are recoverable fro

researcharxiv-cs-cl
31 Jul 2026
Model Releases

LayerRAG-Bench: A Cross-Layer Reliability Benchmark for Agentic Retrieval-Augmented Generation

DGX agent

arXiv:2607.27353v1 Announce Type: new Abstract: Agentic retrieval-augmented generation systems can produce answers that appear grounded while failing at the evidence, tool-contract, authorization, or

model-releasesarxiv-cs-cl
31 Jul 2026
Research

LEEPS: Latent-Guided Explore-Exploit Prompt Sampling for Efficient RLVR in Large Language Models

DGX agent

arXiv:2607.28077v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models, but prompt groups with identical rol

researcharxiv-cs-cl
31 Jul 2026
Model Releases

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models

DGX agent

arXiv:2607.28449v1 Announce Type: new Abstract: On-policy distillation (OPD) provides dense token-level supervision from a teacher, but its effectiveness can depend on teacher consistency, meaning tha

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

LLM Self-Correction with DeCRIM: Decompose, Critique, and Refine for Enhanced Following of Instructions with Multiple Constraints

DGX agent

arXiv:2410.06458v2 Announce Type: replace Abstract: Instruction following is a key capability for LLMs. However, recent studies have shown that LLMs often struggle with instructions containing multipl

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

LLM2Vec-Gen: Generative Embeddings from Large Language Models

DGX agent

arXiv:2603.10913v3 Announce Type: replace Abstract: Fine-tuning LLM-based text embedders via contrastive learning maps inputs and outputs into a new representational space, discarding the LLM's output

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

LLMs struggle to simulate human belief updates in controlled environments

DGX agent

arXiv:2607.28347v1 Announce Type: new Abstract: LLMs are increasingly deployed as proxies for human study participants in social science experiments, yet the fidelity of this practice has rarely been

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

Looped Transformers with Source-Centered State Evolution

DGX agent

arXiv:2607.27656v1 Announce Type: cross Abstract: Looped Transformers create a useful train- and test-time compute axis by reusing the same Transformer block over recurrent depth, increasing effective

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

MagicSelector: Joint Optimization for Agent Tool Selection via Counterfactual Decomposition and Progressive Reranking

DGX agent

arXiv:2607.17751v2 Announce Type: cross Abstract: We present MagicSelector, a joint optimization framework integrating Counterfactual task decomposition, Progressive reranking, and Dynamic Top-K, desi

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

Measuring Alignment With Reader Highlights Net of Position and Length

DGX agent

arXiv:2607.27739v1 Announce Type: cross Abstract: Context compression discards most of a document before a language model reads it, and is normally evaluated by downstream task accuracy - which makes

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

MedHallTune: An Instruction-Tuning Benchmark for Mitigating Medical Hallucination in Vision-Language Models

DGX agent

arXiv:2502.20780v2 Announce Type: replace-cross Abstract: The increasing use of vision-language models (VLMs) in healthcare applications presents great challenges related to hallucinations, in which t

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory

DGX agent

arXiv:2607.27919v1 Announce Type: new Abstract: Decoder-only language models entangle long-term memory and reasoning in a single parameter set, making it difficult to scale memory capacity independent

model-releasesarxiv-cs-cl
31 Jul 2026
Agents

MemTxn: A Transaction Boundary for Source-Supported Updates and Complete-State Recovery in Agent Memory

DGX agent

arXiv:2607.27834v1 Announce Type: cross Abstract: Persistent memory lets long-running large language model agents reuse information across sessions and tasks. Yet errors in writable memory can persist

agentsarxiv-cs-cl
31 Jul 2026
Tutorials

MentorCollab: Selective Large-to-Small Inference-Time Guidance for Efficient Reasoning

DGX agent

arXiv:2602.05307v2 Announce Type: replace Abstract: Large reasoning models (LRMs) achieve strong performance by producing long chains of thought, but their inference costs are high and often generate

tutorialsarxiv-cs-cl
31 Jul 2026
Research

Metaphor Tracer: A Theory-Informed Analysis of Hidden States

DGX agent

arXiv:2607.28434v1 Announce Type: cross Abstract: What do a language model's hidden states say about the organization of a single text? From one forward pass, without training, we score every token po

researcharxiv-cs-cl
31 Jul 2026
Model Releases

Models for minimalist RAG: B1ade 335M Embedding and 1B Parameter Small Language Models

DGX agent

arXiv:2607.27506v1 Announce Type: new Abstract: Language and embedding models used in RAG systems are conventionally assumed to require large-scale pretraining and explicit grounding supervision. We p

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

MORFES: A Benchmark for Productive Inflectional Competence in Modern Greek

DGX agent

arXiv:2607.28274v1 Announce Type: new Abstract: Modern Greek is a richly inflected language, yet the language models built for it are evaluated mainly on factual knowledge, and no benchmark is dedicat

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

ORCA-bench: How Ready Are Language Model Agents for Oncall?

DGX agent

arXiv:2607.28545v1 Announce Type: new Abstract: Large language models can write, patch, and search code, but oncall root cause analysis (RCA) demands something different: reasoning over noisy metrics,

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

DGX agent

arXiv:2607.28609v1 Announce Type: cross Abstract: Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Veri

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

PCAP-LM: An LLM-Native Text Representation for TLS Bulk Traffic Analysis

DGX agent

arXiv:2607.28100v1 Announce Type: cross Abstract: Large language models (LLMs) offer powerful reasoning capabilities for network traffic analysis, but standard capture formats and their textual equiva

model-releasesarxiv-cs-cl
31 Jul 2026
Applications

Prompt Chaining in Practice: A Case Study in Automated Scholarly Report Generation

DGX agent

arXiv:2607.27210v1 Announce Type: new Abstract: The exponential growth of scholarly publications requires automated tools for effective information synthesis. However, simple, single-shot prompting me

applicationsarxiv-cs-cl
31 Jul 2026
Research

Prox: Training-Free FFN Activation Sparsity via Approximate Intermediate-Channel Salience in LLMs

DGX agent

arXiv:2607.27591v1 Announce Type: cross Abstract: Feed-forward networks (FFNs) dominate memory traffic and computation in large language model (LLM) inference, making them a primary target for activat

researcharxiv-cs-cl
31 Jul 2026
Research

Reasoning Consensus: Structural Ensembling of LLM Reasoning via Weighted DAG Aggregation

DGX agent

arXiv:2607.27783v1 Announce Type: new Abstract: Large Language Models (LLMs) explore problems through chain-of-thought, but this exploration is buried in unstructured prose. On high-stakes tasks, user

researcharxiv-cs-cl
31 Jul 2026
Research

Recall Before You Rank: Similarity-Guided Top-K Reuse for Efficient Long-Context Attention

DGX agent

arXiv:2607.27692v1 Announce Type: new Abstract: Top-K sparse attention reduces the cost of Softmax and value aggregation by attending to only a small subset of key--value (KV) entries. However, identi

researcharxiv-cs-cl
31 Jul 2026
Safety

ReDiPPO: Reference-Guided Value Calibration and Discrepancy-Aware Token Reweighting for Mathematical Reasoning

DGX agent

arXiv:2607.27631v1 Announce Type: cross Abstract: Reinforcement learning has emerged as an effective paradigm for enhancing the mathematical reasoning capabilities of large language models. Among exis

safetyarxiv-cs-cl
31 Jul 2026
Model Releases

RepBench: Compiling Benchmarks into Capability Representations for Large Language Models

DGX agent

arXiv:2607.28008v1 Announce Type: new Abstract: Representation engineering reads and steers capability directions in large language models, yet methods are typically evaluated on paper-specific synthe

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models

DGX agent

arXiv:2607.28128v1 Announce Type: new Abstract: LLM tutoring poses a measurement problem: can a general-purpose helpfulness rubric distinguish direct answer-giving from pedagogical guidance? We audit

model-releasesarxiv-cs-cl
31 Jul 2026
Research

RRM: Experience-Driven Reflective Retrieval Memory for Long-Horizon Multimodal Reasoning

DGX agent

arXiv:2607.28156v1 Announce Type: new Abstract: Existing multimodal long-term memory agents use external memory to overcome the limited context available for long videos. However, most methods emphasi

researcharxiv-cs-cl
31 Jul 2026
Safety

Safety Verification of Wait-Only Non-Blocking Broadcast Protocols

DGX agent

arXiv:2403.18591v3 Announce Type: replace-cross Abstract: Broadcast protocols are programs designed to be executed by networks of processes. Each process runs the same protocol, and communication betw

safetyarxiv-cs-cl
31 Jul 2026
Model Releases

Same Facts, Different Diagnosis: Measuring and Mitigating Narrative Anchoring in Clinical Language Models

DGX agent

arXiv:2607.27384v1 Announce Type: new Abstract: Large language models used for clinical diagnostic reasoning are sensitive to sociolinguistic register, not just clinical content. We term this failure

model-releasesarxiv-cs-cl
31 Jul 2026
Research

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B

DGX agent

arXiv:2607.28576v1 Announce Type: new Abstract: Methods that make a language model plan, criticise and rewrite its own answer, reflect on mistakes, pick the best of several attempts, or debate with co

researcharxiv-cs-cl
31 Jul 2026
Model Releases

Scaling medical imaging report generation with multimodal reinforcement learning

DGX agent

arXiv:2601.17151v2 Announce Type: replace-cross Abstract: Frontier models have demonstrated remarkable capabilities in understanding and reasoning with natural-language text, but they still exhibit ma

model-releasesarxiv-cs-cl
31 Jul 2026
Agents

SciDataSailor: Deep Scientific Data Exploring

DGX agent

arXiv:2607.28098v1 Announce Type: cross Abstract: Scientific datasets are commonly organized as hierarchical repositories containing heterogeneous and interdependent files, making their inspection, in

agentsarxiv-cs-cl
31 Jul 2026
Research

SciSchema.org: A Multidisciplinary Collection of Schemas for Structured Scientific Process Descriptions

DGX agent

arXiv:2607.27955v1 Announce Type: cross Abstract: Scientific processes are often described in heterogeneous article discourse, with details needed for comparison, reproducibility, reuse, and automatio

researcharxiv-cs-cl
31 Jul 2026
Model Releases

Select or Project? Evaluating Lower-dimensional Vectors for LLM Training Data Explanations

DGX agent

arXiv:2601.16651v3 Announce Type: replace Abstract: Gradient-based methods for instance-based explanation for large language models (LLMs) are hindered by the immense dimensionality of model gradients

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

Selecting Open-Weight Language Models for Zero-Shot Intent Classification: A Systematic Evaluation of 41 Models

DGX agent

arXiv:2607.27421v1 Announce Type: new Abstract: Intent classification is a core component of task-oriented dialogue systems, yet practitioners have limited systematic guidance for selecting deployable

model-releasesarxiv-cs-cl
31 Jul 2026
Research

Semantic-Aligned Structural Abstraction for Multimodal Sentiment Analysis

DGX agent

arXiv:2607.27790v1 Announce Type: new Abstract: Multimodal Sentiment Analysis (MSA) aims to interpret complex human emotions by integrating natural language with non-verbal modalities. Non-verbal moda

researcharxiv-cs-cl
31 Jul 2026
Agents

SkillSmith: Learning to Compose Parametric Skills and Textual Knowledge

DGX agent

arXiv:2607.27497v1 Announce Type: new Abstract: Agentic systems driven by large language models (LLMs) regularly feature two key mechanisms to autonomously solve complex problems: synthesizing text-ba

agentsarxiv-cs-cl
31 Jul 2026
Research

Stage-Replay Divergence Follows the KV Cache: Fixed-Prefix Precision Controls and Bidirectional Cache Transplantation

DGX agent

arXiv:2607.28495v1 Announce Type: cross Abstract: Stage-replay diagnostics reconstruct intermediate token prefixes and treat fresh-prefill continuation as continuation from the decoder state that orig

researcharxiv-cs-cl
31 Jul 2026
Model Releases

Subtract or Replay? Exact Deletion from Language-Model Memory

DGX agent

arXiv:2607.27539v1 Announce Type: cross Abstract: Exact deletion from persistent language-model memory depends on how that memory represents a record. Addressable influence can be removed by algebraic

model-releasesarxiv-cs-cl
31 Jul 2026
Safety

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute

DGX agent

arXiv:2607.28457v1 Announce Type: cross Abstract: Scaling test-time computation can improve language-model reasoning, but uniform budgets waste computation on easy inputs, while verifier-guided refine

safetyarxiv-cs-cl
31 Jul 2026
Model Releases

Sympathetic Framing: Evaluating AI Alignment across Sociodemographic Groups

DGX agent

arXiv:2607.27232v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly shaping how we consume information and form our worldview. This raises concerns beyond bias in AI: do LLMs

model-releasesarxiv-cs-cl
31 Jul 2026
Research

TCA-SIR: Learning Target-Conditioned Abstractions for Scientific Inspiration Retrieval

DGX agent

arXiv:2607.28498v1 Announce Type: cross Abstract: Scientific hypothesis generation for AI for Science typically involves Scientific Inspiration Retrieval (SIR) followed by hypothesis composition. Exis

researcharxiv-cs-cl
31 Jul 2026
Safety

The Confidence Manifold: Geometric Structure of Correctness Representations in Language Models

DGX agent

arXiv:2602.08159v2 Announce Type: replace-cross Abstract: When a language model asserts that 'the capital of Australia is Sydney,' does it know this is wrong? Models assert misconceptions with the sam

safetyarxiv-cs-cl
31 Jul 2026
Research

The MADRS Pipeline: Supporting Depression Assessment in Clinical Trials

DGX agent

arXiv:2607.28190v1 Announce Type: new Abstract: Depression is a major mental disorder for which diagnosis relies primarily on clinical assessments. Automated methods to support its detection via the p

researcharxiv-cs-cl
31 Jul 2026
Model Releases

THGFM: Dual-Branch Temporal Heterogeneous Graph Fusion Model

DGX agent

arXiv:2607.27303v1 Announce Type: cross Abstract: Temporal heterogeneous graphs offer a natural abstraction for dynamic relational systems in which diverse node and relation types co-exist and evolve

model-releasesarxiv-cs-cl
31 Jul 2026
Agents

ThreatForest: Multi-Agent Attack Tree Generation with Pluggable TTP Framework Mapping

DGX agent

arXiv:2607.27528v1 Announce Type: cross Abstract: Threat modeling is essential for secure software development, yet manual analysis of cloud-native architectures is slow and demands scarce security ex

agentsarxiv-cs-cl
31 Jul 2026
← Previous
1…1617181920…160
Next →