AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
25 Jun 2026

To Isolate or to Score? Model-Adaptive Assessment for Cost-Efficient Multi-Agent RAG

AgentsDGX agent

arXiv:2606.25191v1 Announce Type: cross Abstract: Multi-agent document assessment for retrieval-augmented generation is computationally expensive, driving practitioners toward smaller, deployable mode

Toten: A Knowledge-Based System For Structure-Preserving Representation Of Physical Quantities And Technical Notation In Brazilian Portuguese

Model ReleasesDGX agent

arXiv:2606.19626v2 Announce Type: replace-cross Abstract: AI pipelines that reason quantitatively over technical text depend on input where physical quantities, numbers, units, and symbolic expression

Towards Structuring an Arabic-English Machine-Readable Dictionary Using Parsing Expression Grammars

Research

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2606.25231v1 Announce Type: new Abstract: Dictionaries are rich sources of lexical information about words that is required for many applications of natural language processing and human languag

Tracing Target Answers in Poisoned Retrieval Corpora via Token Influence Attribution

ResearchDGX agent

arXiv:2606.25721v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) systems are vulnerable to corpus poisoning attacks that manipulate model outputs through malicious retrieved docu

Uncertainty Quantification for Computer-Use Agents: A Benchmark across Vision-Language Models and GUI Grounding Datasets

Model ReleasesDGX agent

arXiv:2606.25760v1 Announce Type: cross Abstract: Computer-use agents turn vision-language model (VLM) predictions into executable GUI clicks, so reliable uncertainty estimates are essential for rejec

VADAOrchestra: Neurosymbolic Orchestration of Adaptive Reasoning Workflows

AgentsDGX agent

arXiv:2606.22485v2 Announce Type: replace-cross Abstract: Decision-making in real-world settings rarely follows a fixed script. Instead, it unfolds as a dynamic reasoning process in which the appropri

Weave of Formal Thought

ResearchDGX agent

arXiv:2606.25987v1 Announce Type: new Abstract: Large language models (LLMs) attain remarkable surface fluency on code, yet they neither formally guarantee the syntactic validity of their output nor l

What Intermediate Layers Know: Detecting Jailbreaks from Entropy Dynamics

Model ReleasesDGX agent

arXiv:2606.25182v1 Announce Type: new Abstract: Jailbreak attacks reveal a persistent weakness in aligned Large Language Models: carefully crafted prompts can elicit policy-violating responses despite

When Certainty Is an Artifact: Keyword Lexicon Blindness and the (Mis)Measurement of Rhetorical Stance

ResearchDGX agent

arXiv:2606.26062v1 Announce Type: new Abstract: Can a statistically significant, large-effect-size finding in computational social science be entirely an artifact of the measurement instrument? We pre

Why Do Accumulated Transformations Extrapolate?

ResearchDGX agent

arXiv:2606.24975v1 Announce Type: cross Abstract: PaTH Attention showed that replacing RoPE's position-indexed rotations with accumulated data-dependent Householder reflections yields strong length ex

Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It

SafetyDGX agent

arXiv:2606.26027v1 Announce Type: new Abstract: Tool use enables large language models (LLMs) to perform complex tasks, and recent agentic reinforcement learning (RL) methods show promise for enhancin

24 Jun 2026

A Synthetic Reliability-Aware PINN Benchmark for Offshore Wind Turbine Support-Structure Monitoring with Bayesian Inverse Identification

Model ReleasesDGX agent

arXiv:2606.24176v1 Announce Type: new Abstract: Reliable structural health monitoring (SHM) of offshore wind turbine (OWT) support structures requires fast state estimation from sparse measurements. R

AGORA: An Archive-Grounded Benchmark for Agentic Workplace Document Reasoning

Model ReleasesDGX agent

arXiv:2606.24526v1 Announce Type: new Abstract: Large language models are increasingly deployed as agents that reason over documents rather than answer from parametric knowledge. We study archive-grou

An LLM-based Two-Stage Transformer Framework for Cross-Domain Bearing Fault Diagnosis with Limited Data

SafetyDGX agent

arXiv:2606.24459v1 Announce Type: cross Abstract: Bearing fault diagnosis faces critical challenges when dataset heterogeneity, operating condition variations, and limited labeled data occur simultane

Are We Ready For An Agent-Native Memory System?

Model ReleasesDGX agent

arXiv:2606.24775v1 Announce Type: new Abstract: Memory for large language model (LLM) agents has rapidly evolved from simple retrieval-augmented mechanisms into a data management system that supports

Aspect-Based Sentiment Evolution and its Correlation with Review Rounds in Multi-Round Peer Reviews: A Deep Learning Approach

ResearchDGX agent

arXiv:2606.24188v1 Announce Type: new Abstract: Mining sentiment information from the textual content of peer review comments offers valuable insights into the scientific evaluation process. However,

Automatic Part-of-Speech Tagging of Arabic-English Dictionary Senses through WordNet

ResearchDGX agent

arXiv:2606.24359v1 Announce Type: new Abstract: This paper proposed an algorithm for part-of-speech (POS) tagging senses of a bilingual dictionary. The algorithm is applied on the Al-Mawrid Arabic-Eng

AutoSpecNER: A Fine-Grained Named Entity Recognition Dataset for Vehicle Specification Extraction

Model ReleasesDGX agent

arXiv:2606.24387v1 Announce Type: new Abstract: Vehicle advertisements contain rich specification information, but automotive NER resources remain limited. We introduce AutoSpecNER, an expert-annotate

AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression

ResearchDGX agent

arXiv:2606.24286v1 Announce Type: new Abstract: Multimodal Large Language Models have achieved remarkable progress in short-form audio-video understanding, yet long-form audio-video comprehension rema

BehaviorBench: Benchmarking Foundation Models for Behavioral Science Tasks

Model ReleasesDGX agent

arXiv:2606.24162v1 Announce Type: new Abstract: Foundation models have been increasingly applied to behavioral science domains such as psychology, sociology, and economics. While these models show pro

Best Preprocessing Techniques for Sentiment Analysis

ApplicationsDGX agent

arXiv:2606.24055v1 Announce Type: new Abstract: Sentiment analysis in Twitter datasets is important because it enables monitoring public opinion on products and analysis of political and social moveme

Beyond Logprobs: A Multi-Signal Confidence Engine for LLM-Based Document Field Extraction

ResearchDGX agent

arXiv:2606.24420v1 Announce Type: new Abstract: In high-stakes document processing pipelines, including financial reconciliation, compliance verification, and procurement automation, an LLM extraction

Bilevel Data Curation for LLM Fine-tuning: Offline Selection and Online Self-Refining Generation

SafetyDGX agent

arXiv:2511.21056v2 Announce Type: replace-cross Abstract: Supervised fine-tuning (SFT) datasets are critical to the downstream performance of large language models, yet they often contain low-quality

Business as Rulesual: A Benchmark and Framework for Business Rule Flow Modeling with LLMs

Model ReleasesDGX agent

arXiv:2505.18542v4 Announce Type: replace Abstract: Extracting structured procedural knowledge from unstructured business documents is a critical yet unresolved bottleneck in process automation. While

CANDLE: Character-level Arabic Noise Deduplication using Lightweight Encoder

Model ReleasesDGX agent

arXiv:2606.24758v1 Announce Type: new Abstract: Handling repeated characters in text can be tricky, since they can represent either the correct spelling of a word or informal character elongation ofte

CN-NewsTTS Bench: a target-level automatic benchmark for raw-input Chinese news TTS pronunciation

Model ReleasesDGX agent

arXiv:2606.24714v1 Announce Type: new Abstract: Chinese news text contains dense written forms such as scores, hyphenated model names, ranges, unit symbols, percentages, English abbreviations, and mix

ComputeFHE: A Privacy-Preserving General-Purpose Computation Library

ResearchDGX agent

arXiv:2606.24379v1 Announce Type: cross Abstract: Fully Homomorphic Encryption (FHE) enables computations to be performed directly on encrypted data while preserving data confidentiality. However, its

CORE-BREW: LLR-Based Soft Decoding for Robust Multi-Bit LLM Watermarking

ResearchDGX agent

arXiv:2606.24163v1 Announce Type: cross Abstract: Reliable provenance for LLM outputs requires multi-bit watermarks that remain robust under editing while maintaining strict false-positive control. Ex

Cross-Lingual Exploration for Parametric Knowledge

Local AiDGX agent

arXiv:2606.24579v1 Announce Type: new Abstract: Parametric knowledge in Large Language Models is not equally accessible across languages. As a result, standard inference techniques often struggle to s

Decoherence as Defence and the Magnitude of Noise Regularisation: A Rigorous N -Qubit Theory of Stochastic Quantum Neural Networks for Adversarially Robust Network Intrusion Detection

Local AiDGX agent

arXiv:2606.24219v1 Announce Type: new Abstract: Stochastic quantum neural networks (SQNNs) encode neuronal activations as qubits, synaptic topology as entanglement, and neural noise through a Lindblad

Dialogue to Discovery: Attribute-Aware Preference Elicitation for Conversational Product Search Assistants

ResearchDGX agent

arXiv:2606.24194v1 Announce Type: cross Abstract: Conversational product search assistants offer a more expressive, natural, and interactive alternative to traditional keyword-based product search. Wi

Do LLM Attribution Metrics Transfer? Auditing Retrieval-Augmented Generation Evaluation Across Datasets and Constructs

ResearchDGX agent

arXiv:2606.23915v1 Announce Type: new Abstract: Practice often treats automatic metrics for attribution in LLM retrieval-augmented generation as interchangeable. We audit eight automatic scorers -- le

Does My Embedding Reflect That A = B? Evaluating Mathematical Equivalence in Embedding Models

ResearchDGX agent

arXiv:2606.23959v1 Announce Type: new Abstract: Because mathematics is highly abstract, a single statement can take very different forms depending on what subfield it is framed in. There are many exam

DREAM: Dense Retrieval Embeddings via Autoregressive Modeling

ResearchDGX agent

arXiv:2606.24667v1 Announce Type: new Abstract: Dense retrieval embedding models are a fundamental component of modern retrieval-based AI systems. Most dense retrievers are trained with contrastive ob

ErrorLLM: Modeling SQL Errors for Text-to-SQL Refinement

ResearchDGX agent

arXiv:2603.03742v2 Announce Type: replace Abstract: Despite the remarkable performance of large language models (LLMs) in text-to-SQL (SQL generation), correctly producing SQL queries remains challeng

ESBMC-GraphPLC: Formal Verification of Graphical PLCopen XML Ladder Diagram Programs Using SMT-Based Model Checking

Local AiDGX agent

arXiv:2606.18941v3 Announce Type: replace-cross Abstract: PLCopen XML defines two encoding formats for IEC 61131-3 Ladder Diagram programs: a textual encoding using elements, and a graphical encoding

ESBMC-PLC+: A Unified IEC~61131-3 Formal Verification Framework as a PLCverif Successor

Model ReleasesDGX agent

arXiv:2606.23870v1 Announce Type: cross Abstract: PLCverif is the most mature open-source platform for PLC formal verification, developed at CERN and in production use since 2019. Yet it has two funda

Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning

SafetyDGX agent

arXiv:2606.24428v1 Announce Type: new Abstract: Experience-driven self-evolution is critical for large language model (LLM) agents to improve through open-world interaction. However, existing experien

EvidenceLens: A Claim-Evidence Matrix for Auditing Financial Question Answering

SafetyDGX agent

arXiv:2606.23724v1 Announce Type: cross Abstract: Large language models are increasingly used to answer questions over annual reports, earnings decks, and analyst notes, yet their outputs remain diffi

EXPO-SQL: Execution-based Clause-level Policy Optimization for Text-to-SQL

SafetyDGX agent

arXiv:2606.23693v1 Announce Type: new Abstract: Text-to-SQL enables users to query databases using natural language by generating executable SQL queries. Recent methods have increasingly adopted Large

Few shot chain-of-thought driven reasoning to prompt LLMs for open ended medical question answering

ResearchDGX agent

arXiv:2403.04890v4 Announce Type: replace Abstract: In this paper, we propose a modified version of the MedQA-USMLE dataset, named MEDQA-OPEN, which contains open-ended medical questions without optio

Ground Then Rank: Revisiting Knowledge-Based VQA with Training-Free Entity Identification

ResearchDGX agent

arXiv:2606.23881v1 Announce Type: new Abstract: Knowledge-Based Visual Question Answering (KB-VQA) requires grounding visual queries to external knowledge beyond directly observable content in images.

Harmonic: Hierarchical State Space Models for Efficient Long-Context Language Modeling

Model ReleasesDGX agent

arXiv:2606.24650v1 Announce Type: new Abstract: We present Harmonic, a hierarchical state space model (SSM) for language modeling. The architecture stacks three recurrent levels at progressively slowe

Holistic Data Scheduler for LLM Pre-training via Multi-Objective Reinforcement Learning

Model ReleasesDGX agent

arXiv:2606.24133v1 Announce Type: cross Abstract: The composition of training data, governed by the diversity of sources and their mixing strategy, is a cornerstone of Large Language Model (LLM) pre-t

How Much Can We Trust LLM Search Agents? Measuring Endorsement Vulnerability to Web Content Manipulation

Model ReleasesDGX agent

arXiv:2606.16821v2 Announce Type: replace Abstract: Large language model (LLM)-based search agents synthesize open-web content into actionable recommendations on behalf of users, creating a risk that

Knowledge-Graph Grounding Helps LLMs Only for Out-of-Training Knowledge: A Controlled Study on Clinical Question Answering

Model ReleasesDGX agent

arXiv:2606.22419v2 Announce Type: replace Abstract: A recent Nature Medicine study reports that general-purpose frontier LLMs outperform specialized retrieval-augmented clinical tools on medical bench

L3Cube-MahaPOS: A Marathi Part-of-Speech Tagging Dataset and BERT Models

Model ReleasesDGX agent

arXiv:2606.24825v1 Announce Type: new Abstract: Part-of-Speech (POS) tagging is a foundational NLP task underpinning machine translation, information extraction, and syntactic parsing. Despite Marathi

LangMAP: A Language-Adaptive Approach to Tokenization

Model ReleasesDGX agent

arXiv:2606.23566v2 Announce Type: replace Abstract: Language-specific tokenizers improve tokenization quality and the downstream performance of models on those languages. However, using such a tokeniz

Layer-wise Probing of wav2vec 2.0 and Whisper for Consonant Cluster Reduction in African American English

ResearchDGX agent

arXiv:2606.23948v1 Announce Type: new Abstract: Self-supervised and supervised speech models are increasingly used to investigate which linguistic information their internal representations encode, an

Less is More: Quality-Aware Training Data Selection for Scientific Summarization

SafetyDGX agent

arXiv:2606.24828v1 Announce Type: new Abstract: Scientific long-document summarization datasets commonly treat author-written abstracts as gold reference summaries, although their quality and alignmen

Light-weight Pronunciation Assessment via Discrete Speech Token Surprisal

SafetyDGX agent

arXiv:2606.19910v2 Announce Type: replace Abstract: Training automated pronunciation assessment often relies on labeled learner errors or non-native corpora that are costly to collect. We propose a li

Measuring User's Mental Models of Speech Translation in Human-AI Collaboration

ResearchDGX agent

arXiv:2606.24644v1 Announce Type: new Abstract: Millions of people use machine translation (MT) tools daily, yet little is known about their perception of what systems can and cannot do. This paper st

MedBench v5: A Dynamic, Process-Oriented, and Hallucination-Aware Benchmark for Clinical Multimodal Models

Model ReleasesDGX agent

arXiv:2606.24155v1 Announce Type: new Abstract: Existing medical AI benchmarks lack process visibility, atomic skill evaluation, and integrated hallucination detection. We introduce MedBench v5, a red

Meet UD_Czech-PDTC: A Large and Genre-Rich Treebank in Universal Dependencies

ResearchDGX agent

arXiv:2606.24337v1 Announce Type: new Abstract: Czech has been part of Universal Dependencies since its first release in 2015. It has also been one of the best represented languages, with the Prague D

MEMPROBE: Probing Long-Term Agent Memory via Hidden User-State Recovery

Model ReleasesDGX agent

arXiv:2606.24595v1 Announce Type: new Abstract: Long-term memory promises LLM agents that grow more capable across sessions, maintaining an accurate, evolving understanding of the user that interactio

MERGE: Minimal Expression-Replacement GEneralization Test for Natural Language Inference

ResearchDGX agent

arXiv:2510.24295v2 Announce Type: replace Abstract: As many benchmarks have become saturated, it has become increasingly important to create new datasets that evaluate the generalization capacity of c

ModTGCN: Modularity-aware Graph Neural Networks for Text Classification

ResearchDGX agent

arXiv:2606.23694v1 Announce Type: new Abstract: Graph-based text classification models typically rely on local neighborhood aggregation and overlook global community structure, despite semantic docume

MorfFlex: Handling Rich Morphology

ResearchDGX agent

arXiv:2606.24366v1 Announce Type: new Abstract: We present MorfFlex, a morphological dictionary architecture suitable for languages with extensive regularity in both inflection and derivation. As the

NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers?

Model ReleasesDGX agent

arXiv:2606.24530v1 Announce Type: new Abstract: We introduce NatureBench, a cross-discipline benchmark of 90 tasks distilled from peer-reviewed Nature-family publications, designed to evaluate whether

ParaPairAudioBench: Paralinguistic Pairwise Audio Benchmark for LALM-as-a-Judge

Model ReleasesDGX agent

arXiv:2606.24648v1 Announce Type: cross Abstract: Large Audio-Language Models (LALMs) have been widely used as judge models for the automatic evaluation of generated speech. However, prior approaches

← Previous
1…3536373839…129
Next →