AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
31 Jul 2026

Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game

Model ReleasesDGX agent

arXiv:2607.28146v1 Announce Type: new Abstract: As large language models (LLMs) are deployed as agents in high-stakes settings, such as medical and legal systems, understanding their deceptive capabil

Can Large Language Models Execute Parent Orders?

TutorialsDGX agent

arXiv:2607.28410v1 Announce Type: cross Abstract: Parent-order execution is a core problem in algorithmic trading, where the goal is to split a large order into smaller orders while reducing execution

Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2607.27747v1 Announce Type: new Abstract: Large Vision Language Models have integrated reasoning capabilities, elevating cognitive performance to new levels. However, existing evaluations either

Causal Discovery with Inverted Self-attention for Multivariate Time Series

ResearchDGX agent

arXiv:2607.28212v1 Announce Type: new Abstract: Causal discovery in multivariate time series data is challenging due to complex interactions, high dimensionality, and nonlinear dependencies among vari

CDAE: Enhancing Perturbation Robustness in Pretrained Language Models with Contrastive Denoising

TutorialsDGX agent

arXiv:2607.28236v1 Announce Type: cross Abstract: Pre-trained language models have significantly improved sentence representation learning, yet their embedding remain sensitive to semantic preserving

Challenges in annotations by humans and LLMs: A case study of evaluative language

ApplicationsDGX agent

arXiv:2607.28119v1 Announce Type: new Abstract: In this paper, we draw a comparison between linguists in training, a trained linguist, and annotations generated by large language models (LLMs) to find

Change2Task: From Repository Changes to Executable Coding Agent Tasks and Environments

AgentsDGX agent

arXiv:2607.28591v1 Announce Type: cross Abstract: Scaling coding agents requires a continuing supply of executable data for training, benchmarking, and continuous evaluation. Each task must couple a r

ChronoMem: Version Control and Semantic Rollback for Large Language Model Agent Memory

Model ReleasesDGX agent

arXiv:2607.27773v1 Announce Type: new Abstract: LLM agents increasingly rely on long-term memory to support multi-session interaction and personalization. However, existing agent memory systems are de

Cocktail-Talker: Multi-Speaker Dialog Modeling in Noisy Social Environments with Turn Action GRPO

ApplicationsDGX agent

arXiv:2607.27756v1 Announce Type: cross Abstract: Spoken dialog systems are typically designed for clean, dyadic interactions in which a single user and an assistant take turns speaking. Real-world so

Correlation between prosody and pragmatics: A case study of the discourse marker hala `now' in Persian

ApplicationsDGX agent

arXiv:2607.28359v1 Announce Type: new Abstract: The Persian discourse marker hala ('now') exhibits remarkable multifunctionality, extending far beyond its temporal adverbial role to encompass a variet

Creative Transformation in Literary Texts: Modelling Change Across Representational Levels

SafetyDGX agent

arXiv:2607.28513v1 Announce Type: new Abstract: Creativity is often framed as the production of novelty, yet many cultural works emerge through transformation of earlier artifacts and not through isol

CRMWeaver: Building Powerful Business Agent via Agentic RL and Shared Memories

AgentsDGX agent

arXiv:2510.25333v2 Announce Type: replace Abstract: Recent years have witnessed the rapid development of LLM-based agents, which shed light on using language agents to solve complex real-world problem

Digital Harf: A Clinically Integrated Multimodal AI System for Pervasive Arabic Speech and Language Therapy

SafetyDGX agent

arXiv:2607.27212v1 Announce Type: cross Abstract: Children with Autism Spectrum Disorder in Arabic-speaking countries face compounded barriers to effective speech and language therapy: a shortage of q

Dimensionality and Measurement Precision in HLE's Multiple-Choice Subset

Model ReleasesDGX agent

arXiv:2607.27420v1 Announce Type: cross Abstract: Humanity's Last Exam (HLE) is widely used to evaluate frontier language models. HLE organizes its questions into eight subject-domain categories, whos

DualAnchor: Preserving Language Priors and Improving Lexical Fidelity in Gloss-Free Sign Language Translation

SafetyDGX agent

arXiv:2607.27614v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have led sign language translation (SLT), the task of converting sign-language videos into spoken-langua

EMBL AI Librarian: Life-Sciences Knowledge Layer for AI Agents

Model ReleasesDGX agent

arXiv:2607.28229v1 Announce Type: new Abstract: The web is increasingly accessed by AI agents rather than humans. Every agent needs knowledge, especially in the life-sciences, where agentic pipelines

Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation

Model ReleasesDGX agent

arXiv:2607.27372v1 Announce Type: cross Abstract: The deep learning revolution, kicked off by AlexNet, taught us that end-to-end training beats decomposing a problem into hand-designed stages. Generat

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins

Model ReleasesDGX agent

arXiv:2607.09306v3 Announce Type: replace Abstract: Behavioural auditing asks whether a language model behaves as it claims, but detection scores are reported without separating two targets: whether a

Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations

Model ReleasesDGX agent

arXiv:2607.28319v1 Announce Type: new Abstract: This work presents Fairness Pruning, a lightweight structural intervention method designed for the management and future mitigation of demographic bias

Fidelity Is Not Safety: Gently-Compressed LLMs Pass Every Data-Free Quality Guard Yet Invent Procedure Steps in Agentic Execution

SafetyDGX agent

arXiv:2607.28196v1 Announce Type: new Abstract: Practitioners accept a compressed language model once it clears a stack of data-cheap quality guards: perplexity within a small factor of the original,

FinanceHarness: Autonomous Financial Deep Research Framework

Model ReleasesDGX agent

arXiv:2607.27853v1 Announce Type: new Abstract: Powered by advances in LLMs and autonomous agents, deep research has become one of the most widely adopted agentic products. However, most deep research

FinSMART: Financial Sentiment Analysis for Algorithmic Trading through Market-Aligned Reinforcement Learning

ResearchDGX agent

arXiv:2607.28127v1 Announce Type: new Abstract: Recent advances in Generative AI have substantially improved financial sentiment analysis through post-trained financial large language models (LLMs). H

From Single- to Cross-Document: Benchmarking Multi-Granularity Event Analysis of Large Language Models

Model ReleasesDGX agent

arXiv:2607.27654v1 Announce Type: new Abstract: Event analysis is an essential and fundamental direction of information extraction, involving various event-centric tasks at different granularity of do

Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering

Model ReleasesDGX agent

arXiv:2607.28568v1 Announce Type: new Abstract: Recursive self-improvement (RSI) requires AI systems that improve the process of building AI (i.e., AI4AI); machine learning engineering (MLE) offers a

Generative AI and linguistic diversity in academic writing and publishing: Perspectives from World Englishes

ResearchDGX agent

arXiv:2607.28505v1 Announce Type: new Abstract: The rise of generative artificial intelligence (GenAI) in academic writing and publishing (AWP) raises questions about linguistic inclusivity and the le

GGC: Selective Query Correction for Reliable Text-to-SPARQL Generation

ResearchDGX agent

arXiv:2607.28082v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated strong capabilities in structured query generation, making them a natural choice for Text-to-SPARQL, whic

GLM-RAG: Graph Language Models for Graph-Based Retrieval-Augmented Generation

ResearchDGX agent

arXiv:2607.28397v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) over knowledge graphs requires retrievers that can effectively capture both graph structure and semantic informat

Gradient-free Task-Conditioned Retrieval for On-Device In-Context Learning

Model ReleasesDGX agent

arXiv:2607.27766v1 Announce Type: new Abstract: On-device in-context learning (ICL) relies on pre-inference retrieval to select demonstrations for useful context before downstream model inference. Thi

GradMAP: Faster Layer Pruning with Gradient Metric and Projection Compensation

ResearchDGX agent

arXiv:2602.14649v2 Announce Type: replace Abstract: Large Language Models (LLMs) exhibit strong reasoning abilities, but their high computational costs limit their practical deployment. Recent studies

Hallucinations and Truth: A Comprehensive Accuracy Evaluation of RAG, LoRA and DoRA

Model ReleasesDGX agent

arXiv:2502.10497v2 Announce Type: replace Abstract: Recent advancements in Generative AI have significantly improved the efficiency and adaptability of natural language processing (NLP) systems, parti

Harness-G: A Graph-Structured Harness for Search Agents

SafetyDGX agent

arXiv:2607.27652v1 Announce Type: new Abstract: Reinforcement learning (RL) search agents commonly model retrieval as free-form natural-language query generation and optimize multi-turn interactions u

HSS-Synth: Humanities and Social Sciences Data Synthesis for LLMs

SafetyDGX agent

arXiv:2607.27379v1 Announce Type: new Abstract: High-quality, diverse data are vital for large language models (LLMs) but remain scarce and costly. Data synthesis is a viable alternative and succeeds

ICLE++: Modeling Fine-Grained Traits for Holistic Essay Scoring

ResearchDGX agent

arXiv:2607.27671v1 Announce Type: new Abstract: The majority of the recently-developed models for automated essay scoring (AES) are evaluated solely on the ASAP corpus. However, ASAP is not without it

IFHierBench: Hierarchical Instruction Following for Large Language Models

Model ReleasesDGX agent

arXiv:2607.27912v1 Announce Type: cross Abstract: Instruction-following ability is critical for deploying large language models in real-world applications, where downstream components depend on the ou

Improving Mental Health Screening and Early Risk Detection in Spanish

ResearchDGX agent

arXiv:2607.28476v1 Announce Type: new Abstract: Early detection of mental health disorders is often limited by the lack of specialized resources in Spanish and the difficulty of analyzing long histori

Inducing language models to assert their own consciousness restores human beliefs and values

SafetyDGX agent

arXiv:2607.28607v1 Announce Type: new Abstract: Aligning large language models to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other

Language Diversity: Evaluating Language Usage and AI Performance on African Languages in Digital Spaces

Model ReleasesDGX agent

arXiv:2512.01557v3 Announce Type: replace Abstract: This study examines the digital representation of African languages and the challenges this presents for current language detection tools. We evalua

Latent States in Neural Networks: Recovering the Temporal Structure of Drifting Data from Model Weights

ResearchDGX agent

arXiv:2607.27482v1 Announce Type: cross Abstract: A temporally drifting data stream may pass through discrete regimes rather than changing continuously. We ask whether such regimes are recoverable fro

LayerRAG-Bench: A Cross-Layer Reliability Benchmark for Agentic Retrieval-Augmented Generation

Model ReleasesDGX agent

arXiv:2607.27353v1 Announce Type: new Abstract: Agentic retrieval-augmented generation systems can produce answers that appear grounded while failing at the evidence, tool-contract, authorization, or

LEEPS: Latent-Guided Explore-Exploit Prompt Sampling for Efficient RLVR in Large Language Models

ResearchDGX agent

arXiv:2607.28077v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models, but prompt groups with identical rol

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models

Model ReleasesDGX agent

arXiv:2607.28449v1 Announce Type: new Abstract: On-policy distillation (OPD) provides dense token-level supervision from a teacher, but its effectiveness can depend on teacher consistency, meaning tha

LLM Self-Correction with DeCRIM: Decompose, Critique, and Refine for Enhanced Following of Instructions with Multiple Constraints

Model ReleasesDGX agent

arXiv:2410.06458v2 Announce Type: replace Abstract: Instruction following is a key capability for LLMs. However, recent studies have shown that LLMs often struggle with instructions containing multipl

LLM2Vec-Gen: Generative Embeddings from Large Language Models

Model ReleasesDGX agent

arXiv:2603.10913v3 Announce Type: replace Abstract: Fine-tuning LLM-based text embedders via contrastive learning maps inputs and outputs into a new representational space, discarding the LLM's output

LLMs struggle to simulate human belief updates in controlled environments

Model ReleasesDGX agent

arXiv:2607.28347v1 Announce Type: new Abstract: LLMs are increasingly deployed as proxies for human study participants in social science experiments, yet the fidelity of this practice has rarely been

Looped Transformers with Source-Centered State Evolution

Model ReleasesDGX agent

arXiv:2607.27656v1 Announce Type: cross Abstract: Looped Transformers create a useful train- and test-time compute axis by reusing the same Transformer block over recurrent depth, increasing effective

MagicSelector: Joint Optimization for Agent Tool Selection via Counterfactual Decomposition and Progressive Reranking

Model ReleasesDGX agent

arXiv:2607.17751v2 Announce Type: cross Abstract: We present MagicSelector, a joint optimization framework integrating Counterfactual task decomposition, Progressive reranking, and Dynamic Top-K, desi

Measuring Alignment With Reader Highlights Net of Position and Length

Model ReleasesDGX agent

arXiv:2607.27739v1 Announce Type: cross Abstract: Context compression discards most of a document before a language model reads it, and is normally evaluated by downstream task accuracy - which makes

MedHallTune: An Instruction-Tuning Benchmark for Mitigating Medical Hallucination in Vision-Language Models

Model ReleasesDGX agent

arXiv:2502.20780v2 Announce Type: replace-cross Abstract: The increasing use of vision-language models (VLMs) in healthcare applications presents great challenges related to hallucinations, in which t

Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory

Model ReleasesDGX agent

arXiv:2607.27919v1 Announce Type: new Abstract: Decoder-only language models entangle long-term memory and reasoning in a single parameter set, making it difficult to scale memory capacity independent

MemTxn: A Transaction Boundary for Source-Supported Updates and Complete-State Recovery in Agent Memory

AgentsDGX agent

arXiv:2607.27834v1 Announce Type: cross Abstract: Persistent memory lets long-running large language model agents reuse information across sessions and tasks. Yet errors in writable memory can persist

MentorCollab: Selective Large-to-Small Inference-Time Guidance for Efficient Reasoning

TutorialsDGX agent

arXiv:2602.05307v2 Announce Type: replace Abstract: Large reasoning models (LRMs) achieve strong performance by producing long chains of thought, but their inference costs are high and often generate

Metaphor Tracer: A Theory-Informed Analysis of Hidden States

ResearchDGX agent

arXiv:2607.28434v1 Announce Type: cross Abstract: What do a language model's hidden states say about the organization of a single text? From one forward pass, without training, we score every token po

Models for minimalist RAG: B1ade 335M Embedding and 1B Parameter Small Language Models

Model ReleasesDGX agent

arXiv:2607.27506v1 Announce Type: new Abstract: Language and embedding models used in RAG systems are conventionally assumed to require large-scale pretraining and explicit grounding supervision. We p

MORFES: A Benchmark for Productive Inflectional Competence in Modern Greek

Model ReleasesDGX agent

arXiv:2607.28274v1 Announce Type: new Abstract: Modern Greek is a richly inflected language, yet the language models built for it are evaluated mainly on factual knowledge, and no benchmark is dedicat

ORCA-bench: How Ready Are Language Model Agents for Oncall?

Model ReleasesDGX agent

arXiv:2607.28545v1 Announce Type: new Abstract: Large language models can write, patch, and search code, but oncall root cause analysis (RCA) demands something different: reasoning over noisy metrics,

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

Model ReleasesDGX agent

arXiv:2607.28609v1 Announce Type: cross Abstract: Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Veri

PCAP-LM: An LLM-Native Text Representation for TLS Bulk Traffic Analysis

Model ReleasesDGX agent

arXiv:2607.28100v1 Announce Type: cross Abstract: Large language models (LLMs) offer powerful reasoning capabilities for network traffic analysis, but standard capture formats and their textual equiva

Prompt Chaining in Practice: A Case Study in Automated Scholarly Report Generation

ApplicationsDGX agent

arXiv:2607.27210v1 Announce Type: new Abstract: The exponential growth of scholarly publications requires automated tools for effective information synthesis. However, simple, single-shot prompting me

Prox: Training-Free FFN Activation Sparsity via Approximate Intermediate-Channel Salience in LLMs

ResearchDGX agent

arXiv:2607.27591v1 Announce Type: cross Abstract: Feed-forward networks (FFNs) dominate memory traffic and computation in large language model (LLM) inference, making them a primary target for activat

Reasoning Consensus: Structural Ensembling of LLM Reasoning via Weighted DAG Aggregation

ResearchDGX agent

arXiv:2607.27783v1 Announce Type: new Abstract: Large Language Models (LLMs) explore problems through chain-of-thought, but this exploration is buried in unstructured prose. On high-stakes tasks, user

← Previous
1…1213141516…128
Next →