AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,759 results
Agents

Modeling Distinct Human Interaction in Web Agents

DGX agent

arXiv:2602.17588v3 Announce Type: replace Abstract: Despite rapid progress in autonomous web agents, human involvement remains essential for shaping preferences and correcting agent behavior as tasks

agentsarxiv-cs-cl
2 Jun 2026
Model Releases

Momento: Evaluating Persistent Memory and Reasoning with Multi-Session Agentic Conversations

AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2606.00832v1 Announce Type: new Abstract: Recent advances in agentic AI have enabled agents to complete complex tasks through tool use, reasoning, and multi-step planning. Yet existing benchmark

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

Multi-Agent Computer Use

DGX agent

arXiv:2606.01533v1 Announce Type: cross Abstract: Computer use agents (CUAs) today are primarily deployed as single serial agents. This setup is suboptimal for complex long-horizon tasks that benefit

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

Navigating the Reality Gap: On-Device Continual Adaptation of ASR for Clinical Telephony

DGX agent

arXiv:2512.16401v5 Announce Type: replace Abstract: Automatic Speech Recognition (ASR) can significantly reduce documentation burden in clinical workflows, but standard models degrade sharply in real-

model-releasesarxiv-cs-cl
2 Jun 2026
Safety

NormEval: A Unified Multi-Metric Framework for Evaluating Semantic Fidelity in Text Normalization

DGX agent

arXiv:2511.20409v2 Announce Type: replace Abstract: Text normalization methods such as stemming and lemmatization are fundamental components of NLP pipelines. As new normalization tools are developed

safetyarxiv-cs-cl
2 Jun 2026
Research

Not All Explanations Simulate Equally: Comparing Verbalized Feature Attributions and Self-Generated Rationales

DGX agent

arXiv:2606.01148v1 Announce Type: new Abstract: Natural-language explanations are often treated as a unified interface for understanding model behavior, but different explanation sources may support s

researcharxiv-cs-cl
2 Jun 2026
Agents

Not All Flips Are Conformity: Decomposing Stance Convergence in Multi-Agent LLM Debate

DGX agent

arXiv:2606.00820v1 Announce Type: new Abstract: Multi-agent debate (MAD) is a promising strategy for improving LLM reasoning, but when agents converge on a shared answer, it is unclear whether that co

agentsarxiv-cs-cl
2 Jun 2026
Research

Not What, But How: A Communicative Audit of LLM Response Framing

DGX agent

arXiv:2606.02493v1 Announce Type: new Abstract: Large language models (LLMs) are being increasingly used to answer subjective, information-seeking questions, where users are sensitive to how responses

researcharxiv-cs-cl
2 Jun 2026
Model Releases

OARelatedWork: A Large-Scale Dataset of Related Work Sections with Full-texts from Open Access Sources

DGX agent

arXiv:2405.01930v2 Announce Type: replace Abstract: This paper introduces OARelatedWork: a dataset for related work generation from open-access sources. It is the first large-scale multi-document summ

model-releasesarxiv-cs-cl
2 Jun 2026
Research

OCC-RAG: Optimal Cognitive Core for Faithful Question Answering

DGX agent

arXiv:2606.00683v1 Announce Type: new Abstract: Recent progress in the development of language models has been defined by scale, with each generation absorbing more of the world's knowledge into its w

researcharxiv-cs-cl
2 Jun 2026
Model Releases

OmniOPD: Logit-Free On-Policy Distillation via Speculative Verification

DGX agent

arXiv:2606.01476v1 Announce Type: cross Abstract: On-Policy Distillation (OPD) trains a student model on its own generative trajectories under dense token-level feedback from a stronger teacher, mitig

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

On the Generalization Gap in Self-Evolving Language Model Reasoning

DGX agent

arXiv:2606.01075v1 Announce Type: new Abstract: Recent work suggests that large language models (LLMs) can improve through self-evolution (SE), using supervision signals generated by the model itself.

model-releasesarxiv-cs-cl
2 Jun 2026
Research

On the Salience of Low-Probability Tokens for AI-Generated Text Detection: A Multiscale Uncertainty Perspective

DGX agent

arXiv:2606.02158v1 Announce Type: new Abstract: AI-generated text increasingly blends with human writing, raising practical risks such as misinformation, academic misuse, and corpora contamination. Wh

researcharxiv-cs-cl
2 Jun 2026
Model Releases

On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters

DGX agent

arXiv:2606.02437v1 Announce Type: cross Abstract: Parameter-efficient fine-tuning (PEFT) is usually treated as a cheaper alternative to full fine-tuning. We study a broader role: small trainable adapt

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

OncoReason: Structuring Clinical Reasoning in LLMs for Robust and Interpretable Survival Prediction

DGX agent

arXiv:2510.17532v2 Announce Type: replace Abstract: Predicting cancer treatment outcomes requires models that are both accurate and interpretable, particularly in the presence of heterogeneous clinica

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

PaperVoyager : Building Interactive Web with Visual Language Models

DGX agent

arXiv:2603.22999v3 Announce Type: replace Abstract: Recent advances in visual language models have enabled autonomous agents for complex reasoning, tool use, and document understanding. However, exist

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

Parameter Alignment Mitigates Catastrophic Forgetting in Multilingual Expert Language Models

DGX agent

arXiv:2606.00284v1 Announce Type: new Abstract: While continual pretraining~(CPT) is a practical way to extend large language models to new languages, naive finetuning on targeted data erodes existing

model-releasesarxiv-cs-cl
2 Jun 2026
Applications

Parametric Social Identity Injection and Diversification in Public Opinion Simulation

DGX agent

arXiv:2603.16142v2 Announce Type: replace Abstract: Large language models (LLMs) have recently been adopted as synthetic agents for public opinion simulation, offering a promising alternative to costl

applicationsarxiv-cs-cl
2 Jun 2026
Research

Peacemaker at ATE-IT: Automatic term extraction from Italian text for waste management data using encoder model

DGX agent

arXiv:2606.01469v1 Announce Type: new Abstract: The development of automatic term extraction has become increasingly important in modern technology. Automatic term extraction can be found in virtually

researcharxiv-cs-cl
2 Jun 2026
Research

Phoneme-Level Visual Speech Recognition via Point-Visual Fusion and Language Model Reconstruction

DGX agent

arXiv:2507.18863v2 Announce Type: replace-cross Abstract: Visual Automatic Speech Recognition (V-ASR) is a challenging task that involves interpreting spoken language solely from visual information, s

researcharxiv-cs-cl
2 Jun 2026
Research

PMC-InterCPT: Rethinking Biomedical Interleaved Data for Multimodal Continued Pretraining

DGX agent

arXiv:2606.01049v1 Announce Type: new Abstract: Large-scale biomedical image-text datasets extracted from scientific literature provide valuable resources for medical multimodal model training. These

researcharxiv-cs-cl
2 Jun 2026
Hardware

PortBERT: Navigating the Depths of Portuguese Language Models

DGX agent

arXiv:2606.02100v1 Announce Type: new Abstract: Transformer models dominate modern NLP, but efficient, language-specific models remain scarce. In Portuguese, most focus on scale or accuracy, often neg

hardwarearxiv-cs-cl
2 Jun 2026
Model Releases

Pramana: Fine-Tuning Large Language Models for Epistemic Reasoning through Navya-Nyaya

DGX agent

arXiv:2604.04937v1 Announce Type: cross Abstract: Large language models produce fluent text but struggle with systematic reasoning, often hallucinating confident but unfounded claims. When Apple resea

model-releasesarxiv-cs-cl
2 Jun 2026
Local Ai

ProactiveLLM: Learning Active Interaction for Streaming Large Language Models

DGX agent

arXiv:2606.00523v1 Announce Type: new Abstract: Standard Large Language Models (LLMs) follow a read-then-generate paradigm, causing unnecessary latency and computation. Streaming LLMs alleviate this i

local-aiarxiv-cs-cl
2 Jun 2026
Model Releases

ProtStructQA: A Denotation Threshold in Protein Structural Reasoning

DGX agent

arXiv:2606.00451v1 Announce Type: new Abstract: Protein-language systems are often evaluated by whether they generate plausible biological text, but a structural question has a sharper semantics: it d

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

Quality-Diversity Evolution for Discovering Diverse Vulnerabilities in LLM Safety

DGX agent

arXiv:2606.00801v1 Announce Type: cross Abstract: Current approaches to LLM adversarial testing suffer from coverage gaps: manual red-teaming does not scale, LLM-as-attacker methods exhibit mode colla

model-releasesarxiv-cs-cl
2 Jun 2026
Research

R2-Router: A New Paradigm for LLM Routing with Reasoning

DGX agent

arXiv:2602.02823v2 Announce Type: replace Abstract: As LLMs proliferate with diverse capabilities and costs, LLM routing has emerged by learning to predict each LLM's quality and cost for a given quer

researcharxiv-cs-cl
2 Jun 2026
Tutorials

RCEM: Embedder Equipped with Query Rewriting Skill for Robust Conversational Search in Distributional Shift

DGX agent

arXiv:2606.01697v1 Announce Type: new Abstract: Conversational search has become increasingly important in retrieval-augmented generation (RAG) systems, where users interact with AI assistants through

tutorialsarxiv-cs-cl
2 Jun 2026
Model Releases

RealityTest: How People Probe AI Identity and Whether Models Disclose It

DGX agent

arXiv:2606.00168v1 Announce Type: new Abstract: AI systems are increasingly deployed in conversational settings where users may be uncertain whether they are speaking with a human or an AI. Despite mo

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

Reasmory: 3D Reconstruction as Explicit Memory for VLMs Spatial Reasoning

DGX agent

arXiv:2606.00963v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) exhibit emerging spatial reasoning capabilities, yet they remain unreliable on tasks requiring precise spatial understan

model-releasesarxiv-cs-cl
2 Jun 2026
Safety

Reconsidering Positional Supervision in Masked Diffusion Language Model Training

DGX agent

arXiv:2601.22947v2 Announce Type: replace Abstract: Masked diffusion language models (MDLMs) generate text by unmasking tokens in parallel and have recently emerged as alternatives to autoregressive l

safetyarxiv-cs-cl
2 Jun 2026
Safety

RedDebate: Safer Responses Through Multi-Agent Red Teaming Debates

DGX agent

arXiv:2506.11083v3 Announce Type: replace Abstract: We introduce RedDebate, a novel multi-agent debate framework that provides the foundation for Large Language Models (LLMs) to identify and mitigate

safetyarxiv-cs-cl
2 Jun 2026
Model Releases

RenoBench: A Citation Parsing Benchmark

DGX agent

arXiv:2603.25640v2 Announce Type: replace-cross Abstract: Accurate parsing of citations is necessary for machine-readable scholarly infrastructure. But, despite sustained interest in this problem, exi

model-releasesarxiv-cs-cl
2 Jun 2026
Research

ResMerge: Residual-based Spectral Merging of Large Language Models

DGX agent

arXiv:2606.02252v1 Announce Type: new Abstract: Model merging offers a training-free way to combine multiple post-trained expert models, but merging experts obtained through reinforcement learning (RL

researcharxiv-cs-cl
2 Jun 2026
Model Releases

Resonant Context Anchoring: Decoupling Attention Routing and Signal Gain at Inference Time

DGX agent

arXiv:2606.01923v1 Announce Type: new Abstract: Large Language Models (LLMs) frequently exhibit 'contextual disregard' when faced with input evidence that conflicts with their internal parametric memo

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

Revise, Don't Freeze: Sampler-Matched Training for Self-Correcting Masked Diffusion Language Models

DGX agent

arXiv:2606.01026v1 Announce Type: new Abstract: Masked diffusion language models (MDLMs) re-predict every position at each denoising step, but standard samplers commit tokens once revealed, leaving th

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation

DGX agent

arXiv:2606.01600v1 Announce Type: cross Abstract: Video world models are increasingly used in robotic manipulation, yet existing benchmarks mostly evaluate them under valid, feasible, and safe instruc

model-releasesarxiv-cs-cl
2 Jun 2026
Applications

Robust Asynchronous Planning via Auto-Formalization

DGX agent

arXiv:2606.00981v1 Announce Type: new Abstract: LLMs can plan by either generating action sequences directly as a Planner or translating tasks into domain specific language for an external solver as a

applicationsarxiv-cs-cl
2 Jun 2026
Tutorials

Robust Reasoning via Dynamic Token Selection for Distribution-Aligned Self-Distillation

DGX agent

arXiv:2606.00628v1 Announce Type: new Abstract: Self-distillation improves learning efficiency by rewriting reference answers as training data that better matches the model's own distribution. However

tutorialsarxiv-cs-cl
2 Jun 2026
Research

SALSA: Speech Aware LLM Adaptation via Learned Steering Activation Vectors

DGX agent

arXiv:2606.00460v1 Announce Type: new Abstract: Speech-aware large language models often generalize poorly to out-of-domain settings. We propose SALSA (Speech-Aware LLM Adaptation via Learned Steering

researcharxiv-cs-cl
2 Jun 2026
Model Releases

Same Payload, Different Channel: Measuring Trust Asymmetry in Tool-Using Language Models

DGX agent

arXiv:2606.00566v1 Announce Type: cross Abstract: As language models take on agentic roles that span calling external APIs, reading tool outputs, and acting on instructions embedded in third-party con

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

Sandboxed Coding Agents are Competitive Omni-modal Task Solvers

DGX agent

arXiv:2606.00579v1 Announce Type: new Abstract: As multimodal LLMs increasingly target video and audio, it is often assumed that such tasks require native omnimodal models. We show that this is not al

model-releasesarxiv-cs-cl
2 Jun 2026
Research

SARA: Stress Test Reasoning in Audio Deepfake Detection

DGX agent

arXiv:2601.03615v2 Announce Type: replace Abstract: Audio Language Models (ALMs) offer a promising shift towards explainable audio deepfake detections (ADD), moving beyond extit{black-box} classifiers

researcharxiv-cs-cl
2 Jun 2026
Agents

Scaling Agentic Capabilities via Grounded Interaction Synthesis

DGX agent

arXiv:2606.02001v1 Announce Type: new Abstract: General agentic intelligence hinges on the ability to interact with diverse real-world tools to complete complex tasks, a capability fundamentally tied

agentsarxiv-cs-cl
2 Jun 2026
Agents

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents

DGX agent

arXiv:2602.12984v2 Announce Type: replace Abstract: Scientific reasoning inherently demands integrating sophisticated toolkits to navigate domain-specific knowledge. Yet, current benchmarks largely ov

agentsarxiv-cs-cl
2 Jun 2026
Research

Search-on-Graph: Iterative Informed Navigation for Large Language Model Reasoning on Knowledge Graphs

DGX agent

arXiv:2510.08825v2 Announce Type: replace Abstract: Large language models (LLMs) augmented with knowledge graphs (KGs) offer a promising approach for knowledge-intensive reasoning. Central to this app

researcharxiv-cs-cl
2 Jun 2026
Research

Seeing Through the MiRAGE: Evaluating Multimodal Retrieval Augmented Generation

DGX agent

arXiv:2510.24870v2 Announce Type: replace Abstract: We introduce MiRAGE, an evaluation framework for retrieval-augmented generation (RAG) from multimodal sources. As audiovisual media becomes a preval

researcharxiv-cs-cl
2 Jun 2026
Model Releases

SentGuard: Sentence-Level Streaming Guardrails for Large Language Models

DGX agent

arXiv:2606.02041v1 Announce Type: new Abstract: Large language models increasingly stream long, reasoning-intensive responses in real time, making when to moderate as critical as whether to moderate.

model-releasesarxiv-cs-cl
2 Jun 2026
← Previous
1…6364656667…162
Next →