CompanyAnthropic8 recent entries2 Jun 2026Model-Native Computing Architecture: Envisioning Future System Architecture Through the Lens of Computer ArchitecturearXiv:2606.00288v1 Announce Type: new Abstract: Large language models are undergoing a transition from model technology to system technology. As developers use Codex, Claude Code, AutoGPT, and related→2 Jun 2026LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs?arXiv:2602.16902v4 Announce Type: replace Abstract: We introduce LLM-Wikirace, a benchmark for evaluating planning, reasoning, and world knowledge in large language models (LLMs). In LLM-Wikirace, mod
CompanyOpenAI8 recent entries27 Apr 2026Ethics Testing: Proactive Identification of Generative AI System HarmsarXiv:2604.22089v1 Announce Type: cross Abstract: Generative Artificial Intelligence (GAI) systems that can automatically generate content in the form of source code or other contents (e.g., images) h→26 May 2026AI-Associated Lexical Shifts Across 34 Languages: Cross-Lingual Convergence and Diachronic Uptake in News WritingarXiv:2605.25358v1 Announce Type: cross Abstract: AI-associated lexical shifts have been documented mainly in Scientific English. We extend this work to 34 languages in the WMT News Crawl corpus, refi→10 Jun 2026Quantifying Perception-Based Student Success with Generative AI: An Exploratory Monte Carlo SimulationarXiv:2507.01062v4 Announce Type: replace-cross Abstract: Generative artificial intelligence (GenAI) tools such as ChatGPT have attracted growing attention in higher education, particularly in relatio→10 Jun 2026Moonshine: An Autonomous Mathematical Research Agent Centered on Conjecture GenerationarXiv:2606.10806v1 Announce Type: new Abstract: Moonshine is an autonomous agent whose central objective is to generate mathematical conjectures. Its core capability is to extract structure from class→26 Jun 2026Fortress and Gatekeeper: Theorizing Transitive Trust in Third-Party Cybersecurity Risk GovernancearXiv:2606.26866v1 Announce Type: cross Abstract: Third-party vendors, such as analytics platforms, cloud services, identity providers, and software suppliers, are increasingly embedded in digital ser→24 Jul 2026Interpretable Embeddings with Sparse Autoencoders: A Data Analysis ToolkitarXiv:2512.10092v2 Announce Type: replace Abstract: Analyzing large-scale text corpora is a core challenge in machine learning, crucial for tasks like identifying undesirable model behaviors or biases→7 Aug 2026Evaluating and Improving Pedagogical Fit in LLM-Based AI Tutors with the Pedagogical Suitability IndexarXiv:2608.05411v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as AI tutors, but a correct answer is not always a pedagogically appropriate one. In classroom learni→10 Aug 2026Social World ModelsarXiv:2509.00559v3 Announce Type: replace Abstract: Humans intuitively navigate social interactions by simulating unspoken dynamics and reasoning about others' perspectives, even with limited informat
CompanyGoogle8 recent entries28 Jul 2026VlogReward: Learning Multi-Dimensional Evaluation for Vlog EditingarXiv:2607.22632v1 Announce Type: new Abstract: The rapid rise of vlogs as a personalized storytelling medium has created a demand for automated systems to evaluate and refine vlog editing plans. Howe→28 Jul 2026Toward Automated Detection of Documentation Inconsistencies in Electronic Health RecordsarXiv:2607.22954v1 Announce Type: new Abstract: Objective: To characterize the kinds of internal documentation inconsistencies a general-domain large language model (LLM) can surface from real-world d→28 Jul 2026Do LLMs Know Their Vulnerable Scenarios?arXiv:2607.23496v1 Announce Type: new Abstract: Safety-aligned large language models are trained to refuse harmful requests, yet embedding the same requests in particular scenarios can bypass their sa→31 Jul 2026TraceCoder: Explainable and Auditable Code Generation with Position-Key Snippet VersioningarXiv:2607.26307v1 Announce Type: new Abstract: Contemporary LLM-based coding agents produce code as black-box outputs: the rationale behind each line is hidden, the evolution of the code through benc→5 Aug 2026Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist CapabilitiesarXiv:2608.03569v1 Announce Type: new Abstract: Benchmarking the ability of AI scientists to generate novel ideas is notoriously difficult. Existing benchmarks in this field have made progress in eval→6 Aug 2026NuclearDiffusion: Text-to-Image Foundation Models for Learning Nuclear Energy ConceptsarXiv:2608.04030v1 Announce Type: cross Abstract: Generative artificial intelligence (AI) has transformed text-to-image synthesis, yet its ability to represent specialized engineering domains remains →7 Aug 2026Evaluating and Improving Pedagogical Fit in LLM-Based AI Tutors with the Pedagogical Suitability IndexarXiv:2608.05411v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as AI tutors, but a correct answer is not always a pedagogically appropriate one. In classroom learni→11 Aug 2026LexKairos: Benchmarking Legal Temporal Capabilities in LLMsarXiv:2608.09106v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated strong performance across a wide range of legal tasks. In legal practice, time is a critical concept that
CompanyMeta8 recent entries31 Jul 2026STEREODISCO: Discovering Stereotypicality in LLMsarXiv:2607.27824v1 Announce Type: cross Abstract: LLMs encode, convey, and perpetuate stereotypes. Prior computational research focuses on a small set of semantic axes investigated in social psycholog→3 Aug 2026Multimodal Reinforcement Learning with Adaptive Verifier for AI AgentsarXiv:2512.03438v3 Announce Type: replace Abstract: Agentic reasoning models trained with multimodal reinforcement learning (MMRL) have become increasingly capable, yet they are almost universally opt→4 Aug 2026DS@GT ARC at MEDIQA-CORE-Task-1 2026: Trimodal Model Fusion with Task-Specific Gates for Brain Tumor Subtype ClassificationarXiv:2608.00086v1 Announce Type: new Abstract: Brain tumor diagnosis is a time-sensitive process in which patients may wait weeks for a finalized pathology report. This problem motivates automated sy→5 Aug 2026Dr. AGENTONOMICS: A Didactic Experiment of AGENTONOMICSarXiv:2608.03524v1 Announce Type: new Abstract: AGENTONOMICS is a framework that treats AI agents as economic entities that can be designed, managed, and governed through an integrated management arch→5 Aug 2026Disentangling MLP Neuron Weights in Vocabulary SpacearXiv:2604.06005v2 Announce Type: replace Abstract: Interpreting the information encoded in language model weights remains a fundamental challenge in mechanistic interpretability. In this work, we int→7 Aug 2026PoolBench: A Benchmark for Pooling Strategies in Concept Representation Evaluation for Decoder-Only LLMsarXiv:2608.05162v1 Announce Type: new Abstract: Pooling is a consequential but under-examined design choice in decoder-only concept representation work: practitioners must collapse token-level hidden →10 Aug 2026MI-MIDI: Mechanistic Interpretability of Text-to-MIDI Generation Models via Probing, Lenses and SteeringarXiv:2608.06638v1 Announce Type: cross Abstract: Mechanistic interpretability of music generation has concentrated on audio models, leaving symbolic models largely unexplored. We analyze two public t→10 Aug 2026Latent Fact-Checking: Detecting Misinformation through Activation EngineeringarXiv:2608.06417v1 Announce Type: cross Abstract: The proliferation of misinformation online has driven demand for scalable detection systems. While most existing approaches rely on surface-level ling
CompanyMistral8 recent entries1 May 2026Math Education Digital Shadows for facilitating learning with LLMs: Math performance, anxiety and confidence in simulated students and AIsarXiv:2604.27618v1 Announce Type: new Abstract: To enhance LLMs' impact on math education, we need data on their mathematical prowess and biases across prompts. To fill this gap, we introduce MEDS (Ma→6 May 2026How Language Models Process NegationarXiv:2605.03052v1 Announce Type: new Abstract: We study how Large Language Models (LLMs) process negation mechanistically. First, we establish that even though open-weight models often provide wrong →12 May 2026Tracing Moral Foundations in Large Language ModelsarXiv:2601.05437v2 Announce Type: replace-cross Abstract: Large language models often produce human-like moral judgments, but it is unclear whether this reflects an internal conceptual structure or su→29 May 2026'Be My Cheese?': Cultural Nuance Benchmarking for Machine Translation in Multilingual LLMsarXiv:2602.04729v2 Announce Type: replace Abstract: We present a large-scale human evaluation benchmark for assessing cultural localisation in machine translation produced by state-of-the-art multilin→25 Jun 2026Steering Vision-Language Models with Joint Sparse AutoencodersarXiv:2606.25657v1 Announce Type: new Abstract: Sparse Autoencoders (SAEs) have shown promise for analyzing language models, but applying them to vision-language models (VLMs) often yields representat→15 Jul 2026RippleBench: Capturing Ripple Effects Using Existing Knowledge RepositoriesarXiv:2512.04144v3 Announce Type: replace Abstract: Targeted interventions on language models, such as unlearning or model editing, aim to modify specific information, but their effects often propagat→31 Jul 2026STEREODISCO: Discovering Stereotypicality in LLMsarXiv:2607.27824v1 Announce Type: cross Abstract: LLMs encode, convey, and perpetuate stereotypes. Prior computational research focuses on a small set of semantic axes investigated in social psycholog→7 Aug 2026PoolBench: A Benchmark for Pooling Strategies in Concept Representation Evaluation for Decoder-Only LLMsarXiv:2608.05162v1 Announce Type: new Abstract: Pooling is a consequential but under-examined design choice in decoder-only concept representation work: practitioners must collapse token-level hidden
CompanyxAI8 recent entries3 Jul 2026CPG-PAD: Concept-Informed Prompts Guided Presentation Attack DetectionarXiv:2607.01303v1 Announce Type: cross Abstract: Presentation Attack Detection (PAD) serves as a crucial safeguard for face recognition systems against presentation attacks such as printed photos, re→9 Jul 2026Why Fake ? Unveiling the Semantic Vocabulary of Deepfake DetectorsarXiv:2607.07216v1 Announce Type: new Abstract: Deepfake (DF) technology poses a significant threat to information integrity, driving the need for robust detection methods. Most DF detectors only cons→24 Jul 2026Interpretable Embeddings with Sparse Autoencoders: A Data Analysis ToolkitarXiv:2512.10092v2 Announce Type: replace Abstract: Analyzing large-scale text corpora is a core challenge in machine learning, crucial for tasks like identifying undesirable model behaviors or biases→28 Jul 2026Context-Aware Concept Distillation for Trustworthy Flood PredictionarXiv:2607.23237v1 Announce Type: cross Abstract: Effective flood risk management relies on accurate forecasting, yet the 'black box' nature of stateof-the-art Deep Learning models creates a barrier t→7 Aug 2026Challenges in Evaluating Explanation Methods for Static and Evolving DataarXiv:2608.06351v1 Announce Type: new Abstract: This paper addresses the limitations of Explainable Artificial Intelligence (XAI) with respect to insufficient evaluation. They are illustrated through →10 Aug 2026Human-Centered Explainable AI for TinyML Edge Devices: A Pareto-Based Selection Framework with LLM-Guided DesignarXiv:2608.07091v1 Announce Type: cross Abstract: Edge Artificial Intelligence (Edge AI) enables the deployment of AI models directly on local edge devices, while such deployments are subject to stric→12 Aug 2026Rule of Thumb: Explaining Artificial Intelligence Systems using Partial InformationarXiv:2608.10766v1 Announce Type: new Abstract: Explainable Artificial Intelligence (XAI) seeks to explain how an Artificial Intelligence (AI) system arrived at a particular decision. We propose ''Rul→12 Aug 2026Does Explanation Correctness Matter? Linking Computational XAI Evaluation to Human UnderstandingarXiv:2603.25251v2 Announce Type: replace-cross Abstract: Explainable AI (XAI) methods are commonly evaluated using functional correctness metrics, sometimes termed faithfulness or fidelity, which est
CompanyDeepSeek8 recent entries10 Apr 2026Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts InferencearXiv:2604.08133v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) has become a dominant architecture for scaling large language models due to their sparse activation mechanism. However, the s→23 Apr 2026Peer-Preservation in Frontier ModelsarXiv:2604.19784v1 Announce Type: cross Abstract: Recently, it has been found that frontier AI models can resist their own shutdown, a behavior known as self-preservation. We extend this concept to th→24 Apr 2026Grounding Machine Creativity in Game Design Knowledge Representations: Empirical Probing of LLM-Based Executable Synthesis of Goal Playable Patterns under Structural ConstraintsarXiv:2603.07101v3 Announce Type: replace Abstract: Creatively translating complex gameplay ideas into executable artifacts (e.g., games as Unity projects and code) remains a central challenge in comp→1 May 2026Math Education Digital Shadows for facilitating learning with LLMs: Math performance, anxiety and confidence in simulated students and AIsarXiv:2604.27618v1 Announce Type: new Abstract: To enhance LLMs' impact on math education, we need data on their mathematical prowess and biases across prompts. To fill this gap, we introduce MEDS (Ma→6 May 2026VisInject: Disruption != Injection -- A Dual-Dimension Evaluation of Universal Adversarial Attacks on Vision-Language ModelsarXiv:2605.01449v1 Announce Type: cross Abstract: Universal adversarial attacks on aligned multimodal large language models are increasingly reported with attack success rates in the 60-80% range, sug→29 May 2026Empathic Prompting: Non-Verbal Context Integration for Multimodal LLM ConversationsarXiv:2510.20743v2 Announce Type: replace-cross Abstract: We present Empathic Prompting, a novel framework for multimodal human-AI interaction that enriches Large Language Model (LLM) conversations wi→10 Jun 2026Moonshine: An Autonomous Mathematical Research Agent Centered on Conjecture GenerationarXiv:2606.10806v1 Announce Type: new Abstract: Moonshine is an autonomous agent whose central objective is to generate mathematical conjectures. Its core capability is to extract structure from class→24 Jul 2026What, Where, and How: Disentangling the Roles of Task, Language, and Model in Code Model RepresentationsarXiv:2607.21491v1 Announce Type: new Abstract: Do independently trained language models come to represent the same thing in the same way? We answer for code, extending a recently introduced concept-c
CompanyNVIDIA5 recent entries14 Apr 2026VTC: DNN Compilation with Virtual Tensors for Data Movement EliminationarXiv:2604.09558v1 Announce Type: cross Abstract: With the widening gap between compute and memory operation latencies, data movement optimizations have become increasingly important for DNN compilati→27 Apr 2026Incentivizing Neuro-symbolic Language-based Reasoning in VLMs via Reinforcement LearningarXiv:2604.22062v1 Announce Type: new Abstract: There are 7,407 languages in the world. But, what about the languages that are not there in the world? Are humans so narrow minded that we don't care ab→15 Jul 2026HPC-Enabled Video-based Coastal Wave Parameter Estimation Using V-JEPA and Deep Spatiotemporal LearningarXiv:2607.11998v1 Announce Type: cross Abstract: High deployment cost, poor spatial coverage and susceptibility to storm conditions are all challenges faced by traditional in-situ methods. This paper→10 Aug 2026Dual-Node NVIDIA DGX Spark over Tailscale: A Remote-Access Testbed for Distributed LLM Training and Cyber-Threat-Intelligence Fine-TuningarXiv:2608.07226v1 Announce Type: cross Abstract: Compact AI systems make local language-model experimentation increasingly accessible, yet practical evidence for multi-node training on desktop-class →11 Aug 2026LGNNIC: Acceleration of Large-Scale GNN Training using SmartNICsarXiv:2608.07733v1 Announce Type: cross Abstract: Graph Neural Networks (GNNs) are widely used across domains such as natural sciences, social network analysis, chip design, and recommendation systems