AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,236 results
24 Jul 2026

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers

SafetyDGX agent

arXiv:2607.21010v1 Announce Type: new Abstract: Zero-shot summarization using Large Language Models (LLMs) has significantly advanced the abstractive summarization task by producing coherent and fluen

Refusal-Gated Decoding: Preserving Refusal Behavior Under High-Temperature Sampling

Model ReleasesDGX agent

arXiv:2607.20791v1 Announce Type: new Abstract: High-temperature sampling is one of the primary mechanisms for increasing diversity in LLMs. Recent advances in truncation-based sampling techniques hav

Regulating autonomous and agentic AI

SafetyDGX agent

arXiv:2607.21345v1 Announce Type: new Abstract: Regulating activities where regulatees use autonomous and agentic AI is challenging. Regulatory assumptions about regulatee knowledge and control no lon


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Relative Value Learning

Model ReleasesDGX agent

arXiv:2607.21120v1 Announce Type: cross Abstract: In reinforcement learning, critics typically estimate absolute state values V(s), estimating how good a particular situation is in isolation. However,

Reliability-Aware LLM Alignment from Inconsistent Human Feedback

SafetyDGX agent

arXiv:2607.20515v1 Announce Type: new Abstract: Reinforcement Learning from Human Feedback (RLHF) is critical for aligning Large Language Models (LLMs) with human preferences. However, its efficacy is

ReliableTableQA:How Much Supervision Does Reliability Annotation Need?

ApplicationsDGX agent

arXiv:2607.20537v1 Announce Type: cross Abstract: We introduce ReliableTableQA, a framework for training an LLM to annotate the statistical reliability of tabular QA results, not whether the query is

Representation Robustness Under Executable Reasoning Constraints in Large Language Models for Mathematical Problem Solving

Local AiDGX agent

arXiv:2607.20520v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly evaluated on mathematical problem solving, yet prior work often treats representationally equivalent formu

Representative Sets in Propositional Abduction

ResearchDGX agent

arXiv:2607.21183v1 Announce Type: cross Abstract: The propositional abduction problem is a well-known form of non-monotonic reasoning where we are asked to find an explanation of a given manifestation

Representing Entity Importance in AI Knowledge Systems: A Dual-Signal Framework of Audience Evaluation and Structural Authority

SafetyDGX agent

arXiv:2607.20925v1 Announce Type: new Abstract: AI knowledge systems require representations of entity importance for retrieval, recommendation, evidence selection, and knowledge-intensive reasoning.

Response drift across frontier large language models

ResearchDGX agent

arXiv:2607.20454v1 Announce Type: cross Abstract: All frontier large language models (LLMs) exhibit response drift -- producing outputs that deviate from expert-validated references -- yet the magnitu

Riemannian Deep Learning: Modules, Networks, and Geometries

ResearchDGX agent

arXiv:2607.19305v2 Announce Type: replace-cross Abstract: Deep neural networks on manifold-valued representations have attracted growing interest, but many basic components remain tied to specific man

Robostral Navigate

SafetyDGX agent

arXiv:2607.20785v1 Announce Type: cross Abstract: Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across robot embodiments, and trains efficientl

Robust Critics: Defending LLMs Against Multi-Turn Attacks

SafetyDGX agent

arXiv:2607.20472v1 Announce Type: new Abstract: When a user asks a language model something harmful, is it a genuine attack or a misunderstood but well-meaning question? This ambiguity is one of the c

Routing Subspaces: Auditing Evaluation-to-Deployment Mismatch in Fine-Tuned Language Models

SafetyDGX agent

arXiv:2607.20436v1 Announce Type: cross Abstract: Safety evaluations often assume that behavior observed during testing reflects behavior in ordinary use, but fine-tuning can break this assumption. A

Routing Without Training: Controllable-Ratio LLM Offloading via Reliability Gating

Local AiDGX agent

arXiv:2607.20481v1 Announce Type: new Abstract: Local-cloud collaboration is a practical way to deploy large language models under resource constraints, but existing methods often rely on trained rout

RUMBA: Russian User Memory Benchmark

Model ReleasesDGX agent

arXiv:2607.21447v1 Announce Type: cross Abstract: The ability to handle long-term memory in LLMs is becoming increasingly critical, yet existing benchmarks remain English-centric and rely on aggregate

Safeguards for Speech2Speech LLM-Assistants: A Case Study in Automotive Applications

ApplicationsDGX agent

arXiv:2607.21180v1 Announce Type: new Abstract: Recent advances have introduced speech-to-speech (S2S) conversational assistants capable of producing natural-sounding interactions, including non-verba

SafeStep: AI-powered Travel Assistance for Elderly People with Frailty or Dementia

SafetyDGX agent

arXiv:2607.21156v1 Announce Type: new Abstract: More than a million people in the UK suffer from frailty or dementia, which severely compromise their ability to travel in urban environments. This pape

SalesLoop: Reinforcement Learning from Performance Feedback for Sales Lead Ranking

SafetyDGX agent

arXiv:2607.20655v1 Announce Type: cross Abstract: Lead ranking in Customer Relationship Management (CRM) systems faces a persistent challenge: models achieving high offline accuracy often underperform

Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation

Model ReleasesDGX agent

arXiv:2607.21518v1 Announce Type: new Abstract: Even a current high-capability LLM can appear safer when shown a dangerous objective directly than when other agents transform and relay its direction.

Scaling Closed-Loop Feature Channel Configuration with LLMs

Model ReleasesDGX agent

arXiv:2607.20516v1 Announce Type: cross Abstract: Promising initial results in closed-loop large-language-model-based channel-configuration search demonstrated that neural-network widths can be optimi

Scaling Interpretable Transformers with Parity Bottleneck Layers

Model ReleasesDGX agent

arXiv:2607.20652v1 Announce Type: cross Abstract: Language models are thought to exhibit the phenomenon of superposition, representing many more features than dimensions in their residual streams. Spa

Scaling Up Formal Representation of Clinical Trial Protocols in Ensemble Logic Using LLMs: A Preliminary Study

ApplicationsDGX agent

arXiv:2607.21307v1 Announce Type: cross Abstract: The reliance on unstructured free text for documenting clinical trial protocols creates a significant barrier to automated reasoning, cohort discovery

Scientific exploration, collaboration and labor division in the large language model era

ResearchDGX agent

arXiv:2607.20923v1 Announce Type: cross Abstract: Large language models (LLMs) have rapidly and significantly entered scientific workflows, but it remains unclear how their diffusion is associated wit

SciExplore: Evaluating Autonomous Agents from Scientific Navigation to Information Integration

Model ReleasesDGX agent

arXiv:2607.20926v1 Announce Type: new Abstract: Scientific research involves complex information-seeking and reasoning workflows across heterogeneous sources. However, existing benchmarks primarily em

Self-Evolving Recommendation System: End-To-End Autonomous Model Optimization With LLM Agents

Model ReleasesDGX agent

arXiv:2602.10226v2 Announce Type: replace-cross Abstract: Optimizing large-scale machine learning systems, such as recommendation models for global video platforms, requires navigating a massive hyper

Self-Supervised Bio-Inspired Robotic Trajectory Planning with Obstacle Avoidance

ResearchDGX agent

arXiv:2607.20743v1 Announce Type: cross Abstract: Trajectory planning is a fundamental problem in robotics, requiring the generation of collision-free and efficient trajectories in a potentially compl

Semi-Supervised Text-Attributed Graph Distillation

Model ReleasesDGX agent

arXiv:2607.20477v1 Announce Type: new Abstract: {em Text-Attributed Graphs} (TAGs) have emerged as an expressive data model for integrating graph topology with rich textual semantics. Existing represe

SenCos-GEM: SENet-Calibrated and Law-of-Cosines-Constrained Geometry-Enhanced Molecular Representation for Property Prediction

Model ReleasesDGX agent

arXiv:2607.20551v1 Announce Type: cross Abstract: Effective molecular representation learning is crucial for accurate molecular property prediction. Recently, numerous self-supervised learning (SSL) a

SiGMA: Sign-Guided Merging and Adaptation for Multimodal Continual Instruction Tuning

ResearchDGX agent

arXiv:2607.20511v1 Announce Type: new Abstract: Multimodal Continual Instruction Tuning (MCIT) is crucial for adapting Multimodal Large Language Models (MLLMs) to evolving a sequence of downstream tas

Simple Policy Gradients for Reasoning with Diffusion Language Models

SafetyDGX agent

arXiv:2510.04019v3 Announce Type: replace-cross Abstract: Diffusion large language models (dLLMs) represent a promising alternative to autoregressive LLMs; however, the lack of effective post-training

slang.gr as a Large-Scale Crowdsourced Resource for Non-Standard Greek

SafetyDGX agent

arXiv:2607.21255v1 Announce Type: cross Abstract: Slang is a central component of everyday language, reflecting linguistic creativity, social identity, and cultural change, yet its dy- namic and non-s

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales

Model ReleasesDGX agent

arXiv:2607.20548v1 Announce Type: cross Abstract: Higher-order optimizers such as Muon and SOAP offer faster convergence than AdamW, but their computational cost and numerical stability challenges hav

SonicSampler: Unified Tile-Aware Kernels for LLM Sampling and Speculative Verification

Model ReleasesDGX agent

arXiv:2607.20475v1 Announce Type: new Abstract: Sampling in LLM inference comprises a combinatorial set of logit processing, token selection, and verification operations for speculative decoding. Howe

Source-Prior-Driven Selective Adaptation for Efficient Diffusion Model Finetuning

Model ReleasesDGX agent

arXiv:2607.20913v1 Announce Type: new Abstract: Fine-tuning large diffusion models for new domains or styles involves a trade-off: improving target-specific generation often degrades the pretrained mo

Sparse Concept Channels in Frozen 3D CT Vision Encoders

ResearchDGX agent

arXiv:2607.20993v1 Announce Type: cross Abstract: Large vision-language models are becoming increasingly dominant in 3D medical image interpretation, but we rarely know which internal units encode cli

Spatially Grounded Concept Bottleneck Models for Trustworthy Breast Ultrasound Diagnosis

SafetyDGX agent

arXiv:2607.20691v1 Announce Type: cross Abstract: Concept Bottleneck Models provide interpretable-by-design predictions by mediating diagnosis through human-understandable concepts, but in medical ima

SPORD: A Simulation-Propose-then-OR-Dispose Approach for Supply Chain Planning

HardwareDGX agent

arXiv:2607.21354v1 Announce Type: new Abstract: For years, supply chain planning at e-commerce firms has operated as a collection of isolated projects. Each planning task from static network planning

SR-TTT Does Not Learn Retrieval: A Correction and Mechanistic Post-Mortem of Surprisal-Aware Residual Test-Time Training

TutorialsDGX agent

arXiv:2603.06642v2 Announce Type: replace-cross Abstract: Test-Time Training (TTT) language models replace the KV-cache with fast weights updated during inference, achieving O(1) memory but suffering

StabilityBench: Benchmarking Instability in LLMs

Model ReleasesDGX agent

arXiv:2607.20558v1 Announce Type: cross Abstract: AI Assistants are increasingly deployed in high-stakes settings, such as healthcare or government services. Yet their real-world behavior remains poor

Stochastic Sampling is Epistemically Shallow: The Dimensionality Gap Between Temperature Variation and Model Diversity in LLMs

ResearchDGX agent

arXiv:2607.20464v1 Announce Type: new Abstract: When a language model gives different answers on repeated runs, does that variation reveal what it does not know? Self-consistency turns the variation i

StrideDiffusion: Accelerating Diffusion Models for Time-series Generation

ResearchDGX agent

arXiv:2607.20545v1 Announce Type: new Abstract: Diffusion models have become competitive generators for time series, but their practical use is limited by the large number of sequential denoising step

Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs

Model ReleasesDGX agent

arXiv:2509.14257v3 Announce Type: replace-cross Abstract: Large Language Model agents achieve strong performance on multi-step reasoning and tool-use tasks, but their impressive capabilities typically

Synthetic data generation framework for quality control automation in gravure printing

ApplicationsDGX agent

arXiv:2607.21577v1 Announce Type: cross Abstract: Quality control in printing, particularly in rotogravure printing, still depends on slow, costly, and subjective manual inspection. Automated surface

Synthetic minority data is redundant or invalid: a data-dependent validity theory and a de-biased test

SafetyDGX agent

arXiv:2607.20787v1 Announce Type: cross Abstract: For two decades, the standard remedy for class-imbalanced learning has been to fabricate synthetic minority examples, and the standard evidence of the

TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework

AgentsDGX agent

arXiv:2511.05385v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) utilizes external knowledge to augment Large Language Models' (LLMs) reliability. For flexibility, agenti

Telco-GAIA: Bilingual Benchmark for Agents in Telecom Domain

Model ReleasesDGX agent

arXiv:2607.20510v1 Announce Type: new Abstract: We introduce Telco-GAIA, a bilingual, multi-modal benchmark for evaluating tool-using agents on the data of a real-world telecommunications operator. Te

The Active Ingredient in Muon's Grokking

ResearchDGX agent

arXiv:2607.20512v1 Announce Type: cross Abstract: The Muon optimizer reaches the grokking threshold on modular arithmetic faster than AdamW. Prior work attributes this to 'spectral-norm constraints pl

The Boundaries of Automation: A Theory of Persistent Human Participation

ApplicationsDGX agent

arXiv:2607.21547v1 Announce Type: new Abstract: The rapid progress of AI has intensified the long-standing pursuit of automation: replacing human participation with algorithms wherever possible. Impli

The Devil is in the Spectrum: Mitigating Representation Collapse in LLMs via Topologically Regularized Side-Path

Model ReleasesDGX agent

arXiv:2607.20484v1 Announce Type: new Abstract: Large Language Models (LLMs) are fundamentally limited by representation collapse, a bottleneck that severely degrades long-context performance. We iden

The Geometry of Personality: Activation Steering with Jungian Cognitive Functions

Model ReleasesDGX agent

arXiv:2607.20803v1 Announce Type: cross Abstract: Activation steering enables control and interpretation of LLMs, yet existing work primarily models personality through static trait frameworks such as

The Hidden Footprint: Making Storage a First-Class Metric for LLM Agent Evaluation

Model ReleasesDGX agent

arXiv:2607.11149v3 Announce Type: replace Abstract: LLM agent benchmarks measure task completion, reliability, and inference cost, but not the persistent data an agent run leaves on disk, including lo

The Human-AI Substitution Principle: When will you be replaced by AI in your organization?

ResearchDGX agent

arXiv:2607.20781v1 Announce Type: new Abstract: Artificial Intelligence (AI) is rapidly transforming organizations, raising a fundamental organizational and economic question: when will a human employ

The Storyteller in the Model: Narrative Pattern Inheritance, Escalation Dynamics, and Alignment Governance in LLMs

SafetyDGX agent

arXiv:2607.20449v1 Announce Type: cross Abstract: LLMs are trained predominantly on human-authored text, yet the structural and narrative conventions embedded in that text are rarely examined as a sou

Thinkink: 2D Spatial Ink-native Interaction with LLMs

ResearchDGX agent

arXiv:2607.21468v1 Announce Type: cross Abstract: People often use handwritten notes and sketches to externalize ideas for ideation. To integrate large language models (LLMs) into this practice, we pr

THOR: A Theta-Gamma Hierarchical Oscillatory Reasoning Framework for Multi-hop QA

ResearchDGX agent

arXiv:2607.20459v1 Announce Type: cross Abstract: Multi-hop question answering requires retrieving and integrating evidence from multiple contexts. Despite the rapid progress of current research, mult

Through-the-Earth Magnetic Induction Communication and Networking: A Comprehensive Survey

ResearchDGX agent

arXiv:2510.14854v4 Announce Type: cross Abstract: Magnetic induction (MI) communication (MIC) has emerged as a promising candidate for underground communication networks due to its excellent penetrati

Token Budget Saturation and Mechanistic Early Detection of Reasoning Non-Convergence in Chain-of-Thought Models

Model ReleasesDGX agent

arXiv:2607.21433v1 Announce Type: cross Abstract: Chain-of-thought reasoning models such as DeepSeek-R1-Distill-Qwen-7B exhibit a bimodal convergence pattern: generations either terminate within a tok

TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics

Model ReleasesDGX agent

arXiv:2602.19313v2 Announce Type: replace-cross Abstract: General-purpose robot learning requires dense, instruction-conditioned feedback that can distinguish meaningful task progress from stalled, fa

TOUR: A Trajectory-Level Unlearning Benchmark for Offline Reinforcement Learning

Model ReleasesDGX agent

arXiv:2607.21111v1 Announce Type: cross Abstract: Offline Reinforcement Learning (RL) agents are trained on fixed behavioral trajectories, which makes trajectory-level deletion important when selected

← Previous
1…5758596061…354
Next →