AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlog
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
Safety

Robust Text Watermarking for Large Language Models via Dual Semantic Embeddings

DGX agent

arXiv:2606.31602v1 Announce Type: new Abstract: This work presents Dual-Embedding Watermarking (DEW), a semantic watermarking scheme for large language models (LLMs) that leverages contextual and toke

safetyarxiv-cs-cl
1 Jul 2026
Applications
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Scalable Behaviour Cloning on Browser Using via Skill Distillation

DGX agent

arXiv:2606.32014v1 Announce Type: new Abstract: Internet users collectively perform an enormous range of skilled work through web browsers, from software development and document editing to search, fo

applicationsarxiv-cs-cl
1 Jul 2026
Hardware

SeKV: Resolution-Adaptive KV Cache with Hierarchical Semantic Memory for Long-Context LLM Inference

DGX agent

arXiv:2606.31145v1 Announce Type: new Abstract: Large language models increasingly operate over long contexts, where the KV cache becomes a dominant memory bottleneck: its size grows linearly with seq

hardwarearxiv-cs-cl
1 Jul 2026
Model Releases

SemRF: A Semantic Reference Frame for Residual-Stream Dynamics in Language Models

DGX agent

arXiv:2606.32022v1 Announce Type: cross Abstract: Residual-stream analysis asks how language-model computation evolves across depth, but intermediate decoding requires comparable readout coordinates a

model-releasesarxiv-cs-cl
1 Jul 2026
Model Releases

Signed-Permutation Coordinate Transport for RMSNorm Transformers

DGX agent

arXiv:2606.31963v1 Announce Type: cross Abstract: Modern LLM workflows move coordinate-indexed objects across checkpoints: steering vectors, sparse autoencoders, top-k neuron sets, attribution lists,

model-releasesarxiv-cs-cl
1 Jul 2026
Model Releases

SpikeLogBERT: Energy-Efficient Log Parsing Using Spiking Transformer Networks

DGX agent

arXiv:2606.31781v1 Announce Type: cross Abstract: Log parsing is a fundamental step in automated log analysis, transforming raw system logs into structured event templates for downstream tasks such as

model-releasesarxiv-cs-cl
1 Jul 2026
Research

Symmetry in language statistics shapes the geometry of model representations

DGX agent

arXiv:2602.15029v3 Announce Type: replace-cross Abstract: The internal representations learned by language models consistently exhibit striking geometric structure: calendar months organize into a cir

researcharxiv-cs-cl
1 Jul 2026
Local Ai

TAG-DLM: Diffusion Language Models for Text-Attributed Graph Learning

DGX agent

arXiv:2606.31166v1 Announce Type: new Abstract: Text-attributed graphs (TAGs), where each node carries a natural language description, require models to jointly reason over text and graph topology. Ex

local-aiarxiv-cs-cl
1 Jul 2026
Research

Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability

DGX agent

arXiv:2601.18778v3 Announce Type: replace-cross Abstract: RL methods for scaling large reasoning models stall on datasets with low initial success rates, and thus little training signal. We investigat

researcharxiv-cs-cl
1 Jul 2026
Model Releases

The Bidirectional Process Reward Model

DGX agent

arXiv:2508.01682v3 Announce Type: replace Abstract: Process Reward Models (PRMs), which assign fine-grained scores to intermediate reasoning steps within a solution trajectory, have emerged as a promi

model-releasesarxiv-cs-cl
1 Jul 2026
Model Releases

The Decomposition Is the Fingerprint: Per-Component Identity for Agent Skills

DGX agent

arXiv:2606.31272v1 Announce Type: cross Abstract: AI agents increasingly acquire and execute skills at runtime: bundles of prompt instructions, executable code, and tool declarations fetched from mark

model-releasesarxiv-cs-cl
1 Jul 2026
Model Releases

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action

DGX agent

arXiv:2606.31916v1 Announce Type: new Abstract: Theory of Mind (ToM) benchmarks for Large Language Models (LLMs) typically rely on passive question-answering formats, but the deployment of LLMs in inc

model-releasesarxiv-cs-cl
1 Jul 2026
Applications

Tone-Conditioned Curriculum Learning for Low-Resource Bantu Speech Recognition

DGX agent

arXiv:2606.31642v1 Announce Type: new Abstract: Southern Bantu languages are spoken by over 80 million people, yet current foundation ASR models still produce zero-shot WER above 100%, which limits pr

applicationsarxiv-cs-cl
1 Jul 2026
Model Releases

Truth or Sophistry? LoFa: A Benchmark for LLM Robustness Against Logical Fallacies

DGX agent

arXiv:2606.31039v1 Announce Type: new Abstract: Large Language Models (LLMs) exhibit strong semantic capabilities, yet their resilience to manipulative linguistic patterns such as logical fallacies re

model-releasesarxiv-cs-cl
1 Jul 2026
Research

Usage frequency and application variety of research methods in library and information science: Continuous investigation from 1991 to 2021

DGX agent

arXiv:2606.31081v1 Announce Type: cross Abstract: The present study analyzed over 26,000 research articles published between 1991 and 2021 in twenty-one major LIS (Library and Information Science) jou

researcharxiv-cs-cl
1 Jul 2026
Agents

Using AI Agents to Automate Black-Box Audits of Personalization Algorithms at Scale

DGX agent

arXiv:2606.30801v1 Announce Type: new Abstract: Personalization algorithms determine what content users encounter on online platforms. Auditing these systems is difficult because independent auditors

agentsarxiv-cs-cl
1 Jul 2026
Tutorials

ViTL: Temporal Logic-Guided Zero-Shot Natural Language Navigation via Vision-Language Models

DGX agent

arXiv:2606.30696v1 Announce Type: cross Abstract: Enabling robots to follow natural language commands to complete zero-shot long-horizon tasks remains challenging. It requires extracting implicit temp

tutorialsarxiv-cs-cl
1 Jul 2026
Model Releases

What Counts as an Error? Dual-Reference Benchmarking for Atypical ASR

DGX agent

arXiv:2606.31112v1 Announce Type: new Abstract: ASR systems have been often reported to underperform on atypical speech. An often conflated compounding factor is the existence of two valid transcripti

model-releasesarxiv-cs-cl
1 Jul 2026
Model Releases

What If We Allocate Test-Time Compute Adaptively?

DGX agent

arXiv:2602.01070v5 Announce Type: replace Abstract: Test-time compute scaling allocates inference computation uniformly, uses fixed sampling strategies, and applies verification only for reranking. In

model-releasesarxiv-cs-cl
1 Jul 2026
Research

When Calibration Rankings Reverse: Accuracy-Controlled Evaluation for Fair Comparison of LLMs

DGX agent

arXiv:2606.30814v1 Announce Type: new Abstract: Calibration evaluates whether a model confidence aligns with its empirical accuracy. Existing studies often compare the calibration of different large l

researcharxiv-cs-cl
1 Jul 2026
Model Releases

When the Database Fails: Prompting LLM Dialogue Agents for Safe Recovery in Task-Oriented Dialogue

DGX agent

arXiv:2606.31307v1 Announce Type: new Abstract: Large language models used in task-oriented dialogue often produce fluent but unsafe responses when backend database calls fail, return empty results, o

model-releasesarxiv-cs-cl
1 Jul 2026
Model Releases

A Diagnostic Framework and Multi-Evaluator Audit of Evaluator-Driven Preference Dynamics in Self-Adapting LLM Agents

DGX agent

arXiv:2606.29719v1 Announce Type: cross Abstract: Measurements of proprietary LLM evaluators can become invalid within weeks -- we document one case and provide the diagnostic framework to detect it.

model-releasesarxiv-cs-cl
30 Jun 2026
Research

A French OSCE Dialogue Dataset and Controllable Virtual Patient System for Clinical Training

DGX agent

arXiv:2606.28526v1 Announce Type: new Abstract: The clinical and communication skills of medical students are commonly assessed through Objective Structured Clinical Examinations (OSCEs), which consis

researcharxiv-cs-cl
30 Jun 2026
Applications

A Good Talk Does not Look Like a Summary, It Teaches You! Measuring Takeaways from Paper-to-Video Talks

DGX agent

arXiv:2606.28531v1 Announce Type: cross Abstract: Automatically generated videos from scientific papers are increasingly used for education and research dissemination. However, existing evaluation met

applicationsarxiv-cs-cl
30 Jun 2026
Safety

A3M: Adaptive, Adversarial and Multi-Objective Learning for Strategic Bidding in Repeated Auctions

DGX agent

arXiv:2606.28943v1 Announce Type: new Abstract: Learning to bid in repeated multi-unit auctions with bandit feedback poses a fundamental challenge. Existing methods often rely on rigid explore-then-ex

safetyarxiv-cs-cl
30 Jun 2026
Agents

Adam's Law: Textual Frequency Law on Large Language Models

DGX agent

arXiv:2604.02176v3 Announce Type: replace Abstract: While textual frequency has been validated as relevant to human cognition in reading speed, its relatedness to Large Language Models (LLMs) is seldo

agentsarxiv-cs-cl
30 Jun 2026
Safety

Agentic Tool Use in Large Language Models

DGX agent

arXiv:2604.00835v2 Announce Type: replace Abstract: Large language models are increasingly being deployed as autonomous agents yet their real world effectiveness depends on reliable tools for informat

safetyarxiv-cs-cl
30 Jun 2026
Local Ai

An Information-Geometric Justification for Composite Coherence in Event-Based Narrative Extraction

DGX agent

arXiv:2606.29118v1 Announce Type: cross Abstract: Graph-based narrative extraction relies on a coherence function to score transitions between events, but the coherence metrics in current use are defi

local-aiarxiv-cs-cl
30 Jun 2026
Model Releases

Anisotropy Decides Cosine vs. Rank Metrics for Text Embeddings

DGX agent

arXiv:2606.29571v1 Announce Type: new Abstract: The standard way to compare two text embeddings is cosine similarity. Scattered studies report that a different metric does better, but never pin down t

model-releasesarxiv-cs-cl
30 Jun 2026
Research

AnTenA: Actionable and Explainable Tensor Analysis System with Large Language Models

DGX agent

arXiv:2606.28708v1 Announce Type: new Abstract: Accurately explaining hidden patterns in multi-aspect data has typically been done by leveraging labels and/or accompanying auxiliary metadata. However,

researcharxiv-cs-cl
30 Jun 2026
Safety

Are Humans Evolved Instruction Followers? An Underlying Inductive Bias Enables Rapid Instructed Task Learning

DGX agent

arXiv:2606.29792v1 Announce Type: new Abstract: Human adults can often perform a novel task correctly on the first attempt after only receiving verbal or written instructions. This rapid instructed ta

safetyarxiv-cs-cl
30 Jun 2026
Safety

Are We Measuring Strategy or Phrasing? The Gap Between Surface- and Approach-Level Diversity in LLM Math Reasoning

DGX agent

arXiv:2606.29985v1 Announce Type: new Abstract: Diversity in LLM mathematical reasoning is critical for exploration, but common diversity metrics mostly capture surface-level variation rather than dif

safetyarxiv-cs-cl
30 Jun 2026
Model Releases

Attractor States Emerge in Multi-Turn LLM Conversations

DGX agent

arXiv:2606.30571v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used in open-ended multi-agent settings, but the long-run dynamics of model--model interaction remain po

model-releasesarxiv-cs-cl
30 Jun 2026
Model Releases

AURORA: Asymmetry and Update-Induced Rotation for Robust Hallucination Detection in Large Language Models

DGX agent

arXiv:2606.29545v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities across a wide range of natural language processing tasks. However, their tendency

model-releasesarxiv-cs-cl
30 Jun 2026
Tutorials

Before Thinking, Learn to Decide: Proactive Routing for Efficient Visual Reasoning

DGX agent

arXiv:2606.30217v1 Announce Type: new Abstract: Large multimodal models have achieved strong reasoning on complex visual tasks, but their inference efficiency is often restricted by long chains of tho

tutorialsarxiv-cs-cl
30 Jun 2026
Agents

Benchmarking LLM Agents on Meta-Analysis Articles from Nature Portfolio

DGX agent

arXiv:2606.17041v4 Announce Type: replace Abstract: Meta-analysis is a demanding form of evidence synthesis that combines literature retrieval, PI/ECO-guided study selection, and statistical aggregati

agentsarxiv-cs-cl
30 Jun 2026
Model Releases

Beyond the Mean: Three-Axis Fidelity for Aligning LLM-Based Survey Simulators from Small Pilot Data

DGX agent

arXiv:2606.28963v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to simulate social survey responses, yet their outputs exhibit systematic biases: marginal distributi

model-releasesarxiv-cs-cl
30 Jun 2026
Model Releases

Can LLM-as-a-Judge Reliably Verify Rubrics in Agentic Scenarios?

DGX agent

arXiv:2606.29920v1 Announce Type: new Abstract: Rubric-based scoring has become a widely used paradigm in model evaluation, typically with LLM-as-a-Judge (LaaJ) for rubric scoring. However, the reliab

model-releasesarxiv-cs-cl
30 Jun 2026
Model Releases

Can LLMs Hire Fairly? Racial Bias in Resume Screening

DGX agent

arXiv:2606.28978v1 Announce Type: new Abstract: We audit fourteen mainstream large language models (LLMs) for hiring discrimination using the paired-resume methodology of Kline, Rose, and Walters (202

model-releasesarxiv-cs-cl
30 Jun 2026
Safety

Can LLMs Reliably Self-Report Adversarial Prefills, and How?

DGX agent

arXiv:2606.23671v2 Announce Type: replace Abstract: Prior work shows that large language models (LLMs) exhibit introspective capability on benign tasks. We extend the question to safety contexts and e

safetyarxiv-cs-cl
30 Jun 2026
Safety

Can MLLMs Critique Like Humans? Evaluating Open-Ended Aesthetic Reasoning in Multimodal Large Language Models

DGX agent

arXiv:2606.29689v1 Announce Type: new Abstract: Open-ended aesthetic critique is a challenge for multimodal large language models (MLLMs): unlike multiple-choice aesthetic benchmarks, it has no single

safetyarxiv-cs-cl
30 Jun 2026
Model Releases

Can OCR-VLMs Read Devanagari? A Stress-Test Benchmark and Post-Correction Study

DGX agent

arXiv:2606.29213v1 Announce Type: new Abstract: OCR systems, ranging from classical engines to specialised OCR vision-language models (OCR-VLMs) and frontier multimodal LLMs, report strong results on

model-releasesarxiv-cs-cl
30 Jun 2026
Safety

CaresAI at CT-DEB26: Detecting Dosing Errors In Clinical Trials Using Domain-Specific Transformer Embeddings and Classification Models

DGX agent

arXiv:2606.30236v1 Announce Type: new Abstract: Medication errors, particularly dosing errors in clinical trials (CT), can lead to patient harm, adverse drug events and worse patient outcomes. Dosing

safetyarxiv-cs-cl
30 Jun 2026
Research

Categorize Early, Integrate Late: Divergent Processing Strategies in Automatic Speech Recognition

DGX agent

arXiv:2601.06972v2 Announce Type: replace Abstract: In speech language modeling, two architectures dominate the frontier: the Transformer and the Conformer. However, it remains unknown whether their c

researcharxiv-cs-cl
30 Jun 2026
Applications

Comparing Human and Automatic Recognition of Dutch Dysarthric Continuous Speech: A Case Study

DGX agent

arXiv:2606.30237v1 Announce Type: new Abstract: In our goal to develop personalised dysarthric speech recognition (DSR) models, this study compared the recognition performances of human listeners and

applicationsarxiv-cs-cl
30 Jun 2026
Safety

Complementary RL: Towards Efficient Experience-Driven Agent Learning

DGX agent

arXiv:2603.17621v2 Announce Type: replace-cross Abstract: Reinforcement Learning (RL) has emerged as a powerful paradigm for training LLM-based agents, yet remains limited by low sample efficiency, st

safetyarxiv-cs-cl
30 Jun 2026
Model Releases

Compressed Sensing for Capability Localization in Large Language Models

DGX agent

arXiv:2603.03335v2 Announce Type: replace Abstract: Large language models (LLMs) exhibit a wide range of capabilities, including mathematical reasoning, code generation, and linguistic behaviors. We s

model-releasesarxiv-cs-cl
30 Jun 2026
Research

Conversational Domain Adaptation of IndicTrans2 across 21 Indic Languages via Experience Replay and Model Soups

DGX agent

arXiv:2606.29024v1 Announce Type: new Abstract: IndicTrans2 is the strongest open English to Indic translation system, but like most systems it is trained on general text and tends to sound stiff on c

researcharxiv-cs-cl
30 Jun 2026
← Previous
1…3637383940…161
Next →