CompanyAnthropic8 recent entries5 Aug 2026Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent EvaluationarXiv:2608.03166v1 Announce Type: new Abstract: Role-Playing Language Agents (RPLAs) are increasingly deployed in high-stakes applications such as healthcare assistance, customer support, and educatio→6 Aug 2026One Surrogate to Fool Them All: Universal, Transferable, and Targeted Adversarial Attacks with CLIParXiv:2505.19840v3 Announce Type: replace-cross Abstract: Deep Neural Networks (DNNs) have achieved widespread success yet remain prone to adversarial attacks. Typically, such attacks either involve f
CompanyOpenAI8 recent entries8 Aug 2026You may have been told to watch this video about the OpenAI AI hack. You really should, even if you don't usually care about tech stuff. If …You may have been told to watch this video about the OpenAI AI hack. You really should, even if you don't usually care about tech stuff. If nothing else, click this link to the 18 minutes in & see how→10 Aug 2026We've used GPT-5.6-Cyber extensively in real-world vulnerability research, including work that uncovered previously unknown vulnerabilities …OpenAI announced the release of GPT‑5.6‑Cyber as part of its Cybersecurity Initiative, “Daybreak.” The model is aimed at advanced, authorized security research and testing, helping trusted defenders d→10 Aug 2026Ultra-High Interactivity on NVIDIA GPUs? - TileRT InferenceXThe TileRT InferenceX article (Aug 10 2026) examines whether the TileRT software stack on NVIDIA GPUs can compete with dedicated inference systems such as Cerebras, Groq LPUs and SambaNova for ultra‑h→10 Aug 2026Blast RadiusarXiv:2608.07440v1 Announce Type: new Abstract: Agentic coding faces growing problems of affordability and wasted tokens. We introduce Blast Radius, a predictive memory management layer that estimates→11 Aug 2026Now in preview: The ChatGPT desktop app for Linux. Use ChatGPT, ChatGPT Work, and Codex where you already work and build, with your projects…OpenAI released a preview of its new ChatGPT desktop application for Linux on August 11 2026. The app allows users to run ChatGPT, ChatGPT Work and Codex directly from their desktop, integrating with →11 Aug 2026How to Ask the AI: A User Perspective Survey for Large Language Model PromptingarXiv:2608.07494v1 Announce Type: cross Abstract: AI tools like ChatGPT and DeepSeek, powered by Large Language Models (LLMs), allow users to obtain instant and effective content responses simply by t→12 Aug 2026Rethinking LLM Verification: Evidence Structure, Uncertainty, and Selective RefinementarXiv:2608.10725v1 Announce Type: new Abstract: Large language models (LLMs) often rely on shortcuts rather than systematic reasoning, raising safety concerns in medical applications. Allowing models →12 Aug 2026I believe the demand for compute is going to go up faster than the supply of compute, so the price of compute is going to increase significa…I believe the demand for compute is going to go up faster than the supply of compute, so the price of compute is going to increase significantly in the future, perhaps as much as 10x in the next few y
CompanyGoogle8 recent entries6 Aug 2026Mirendil taps AI Hypercomputer TPUs and GPUs for pre- and post-training applicationsNearly every major AI lab uses Google Cloud infrastructure, including for training of models, inference for agents, and new frontier research. Google Cloud also continues to be the platform of choice →6 Aug 2026Agentic Future Ready With BigQuery: Continually Improving Price-Performance, Zero EffortIn the modern data landscape, query performance tuning and managing system price-performance is challenging, especially as the number of agentic workloads increase. Even for experienced developers and→10 Aug 2026Introducing the Developer Device Platform for agentic mobile app developmentMost enterprises connect with their customers through a device. Whether it’s using a mobile app to order a product, contact customer service, view content, or manage their account, the customer experi→10 Aug 2026Google named a Leader in The Forrester Wave™: AI Platforms, Q3 2026At Google Cloud, we help organizations of all sizes build and operationalize complex agentic workflows with total confidence. By combining world-class AI research with an open, fully integrated AI pla→11 Aug 2026VideoVIBE: A Video-Grounded Diagnostic Benchmark for One-Shot Interactive Website GenerationarXiv:2608.09573v1 Announce Type: new Abstract: Natural-language-driven 'vibe coding' enables the one-shot generation of visually rich and interactive web applications, yet reliable assessment of thei→11 Aug 2026Temporal Misgrounding in Legal RAG: A Versioned-Corpus Benchmark for French Tax LawarXiv:2608.09393v1 Announce Type: cross Abstract: We identify and quantify temporal misgrounding: the systematic retrieval and citation of the currently in-force version of a legal article when the ap→11 Aug 2026LLM within MCP Matters: Measuring Inefficient Resource Utilization Driven by LLMsarXiv:2608.08467v1 Announce Type: new Abstract: The Model Context Protocol (MCP) standardizes how servers expose data and tools to Large Language Models (LLMs). A common server design embeds frequentl→11 Aug 2026Accelerate PostgreSQL migrations using Gemini in Database Migration ServiceImagine this scenario: Your team decides to migrate a core application from an existing commercial database like Oracle or SQL Server to open source PostgreSQL or a fully managed service such as Alloy
CompanyMeta8 recent entries10 Aug 2026Introducing Muse GlimmerIntroducing Muse Glimmer Meta are back in the open weights game! Muse Glimmer is a brand new 30B model under a clean Apache 2.0 license (a step up from the janky Llama licenses of old). They claim to →10 Aug 2026Counterfactual Simulation Training for Chain-of-Thought FaithfulnessarXiv:2602.20710v2 Announce Type: replace Abstract: Inspecting Chain-of-Thought reasoning is among the most common means of understanding why an LLM produced its output. But well-known problems with C→10 Aug 2026Coding isn't yet another application domain -- it's the meta-skill required for AI to automatically develop its own training material, via s…Coding isn't yet another application domain -- it's the meta-skill required for AI to automatically develop its own training material, via symbolic world models. That's how the RSI loop actually kicks→10 Aug 2026Beyond Fluency: A Clinical Benchmark and Anomaly-Enhanced Baseline for Spine MRI Report GenerationarXiv:2608.07117v1 Announce Type: new Abstract: Radiology reporting is time-consuming and subject to inter-rater variability, making automated report generation an attractive clinical application for →11 Aug 2026Hybrid Neural-Classical Correction for Frozen Time Series Foundation Models: A Comprehensive Ablation Study on High-Frequency Stock PredictionarXiv:2608.08825v1 Announce Type: cross Abstract: Foundation models for time series forecasting demonstrate impressive zero-shot generalization but often underperform on specialized domains such as hi→11 Aug 2026How sensitive do we want AI to be? Socio-communicative competencies of large language models in healthcarearXiv:2608.07511v1 Announce Type: cross Abstract: Background. Effective clinical practice relies heavily on the socio-communicative skills of medical professionals. Large language models (LLMs) have b→12 Aug 2026Quantum Coordination Advantages in AI State-Tracking Tasks: Semantic Compilation and Latent MemoryarXiv:2608.11066v1 Announce Type: cross Abstract: We prove inference-time quantum coordination advantages for specified AI state-tracking tasks. A solver compresses semantic history into a future-acce→12 Aug 2026Measuring Semantic Abstractness of SAE Features via NonlocalityarXiv:2608.10537v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) have helped uncover mechanistic explanations for LLM behaviours such as reasoning, jailbreaking etc., via understanding the c
CompanyMistral8 recent entries8 Jul 2026RPAM: A Principled Metric for Evaluating Associations in Language Models with High Predictive Validity in Downstream OutputsarXiv:2607.05679v1 Announce Type: cross Abstract: Language models (LMs) exhibit problematic biases, such as stereotypes. Effectively analyzing and mitigating such biases requires accurate and generali→8 Jul 2026It runs on wheeled, legged, and flying robots and generalizes across sizes, unlocking delivery, logistics, manufacturing, and hospitality. R…Mistral AI has developed a unified robotic control system or framework that operates across multiple robot morphologies—including wheeled, legged, and flying variants—while generalizing across differe→15 Jul 2026Wiki Lint Report — 2026-07-15Automated lint: 26 errors, 6728 warnings, 3 info→19 Jul 2026Wiki Lint Report — 2026-07-19Automated lint: 20 errors, 8743 warnings, 3 info→29 Jul 2026Decision-Level Hijacking: Injecting Cognitive Bias into Large Language Models via Bit-Flip AttacksarXiv:2607.25227v1 Announce Type: cross Abstract: Large Language Models (LLMs) have been widely applied in high-stakes decision-making scenarios such as corporate strategy, and users are increasingly →31 Jul 2026What’s new in AI infrastructure and orchestration this monthAt Google, AI is a soup-to-nuts endeavor. Obviously, we make leading AI models like Gemini and Nano Banana. We incorporate AI into the tools you use every day (think Gmail, BigQuery, AlloyDB, Google C→4 Aug 2026Feed-Forward Steering in Transformer Residual DynamicsarXiv:2608.02071v1 Announce Type: new Abstract: Attention-only dynamical theories model Transformer residual directions as particles aggregating on a sphere. We extend this framework by incorporating →6 Aug 2026nvidia/NVIDIA-Nemotron-Parse-2.0 · Hugging FaceNVIDIA Nemotron Parse 2.0 transforms document images into structured, machine-readable representations with text, layout classes, bounding boxes, and reading-order information. Given a Red, Green, Blu
CompanyxAI8 recent entries28 Jul 2026Evaluating the Impact of Explainable AI on Trust in AI-Assisted Code ReviewarXiv:2607.24601v1 Announce Type: cross Abstract: Background: Large language models (LLMs) are increasingly used to automate code review, but the reasoning behind their decisions remains hard to under→29 Jul 2026On the Design and Evaluation of Human-centered Explainable AI Systems: A Systematic Review and TaxonomyarXiv:2510.12201v2 Announce Type: replace Abstract: As AI becomes more common in everyday living, there is an increasing demand for intelligent systems that are both performant and understandable. Exp→29 Jul 2026Explainable AI for Chronic Kidney Disease Prediction Using Simulated Federated LearningarXiv:2607.25348v1 Announce Type: cross Abstract: Chronic Kidney Disease (CKD), characterized by the gradual loss of kidney function, remains a significant public health challenge. Early detection is →3 Aug 2026A Novel XAI-Enhanced Quantum Adversarial Networks for Velocity Dispersion Modeling in MaNGA GalaxiesarXiv:2510.24598v2 Announce Type: replace Abstract: Current quantum machine learning approaches often face challenges balancing predictive accuracy, robustness, and interpretability. To address this, →5 Aug 2026Explainable AI for the EU Right to Explanation: A Systematic Review of the Law-XAI Translation GaparXiv:2608.02699v1 Announce Type: new Abstract: When algorithms make or influence consequential decisions---about loan eligibility, hiring, or healthcare---EU law grants affected individuals a Right t→10 Aug 2026Human-Centered Explainable AI for TinyML Edge Devices: A Pareto-Based Selection Framework with LLM-Guided DesignarXiv:2608.07091v1 Announce Type: cross Abstract: Edge Artificial Intelligence (Edge AI) enables the deployment of AI models directly on local edge devices, while such deployments are subject to stric→12 Aug 2026Does Explanation Correctness Matter? Linking Computational XAI Evaluation to Human UnderstandingarXiv:2603.25251v2 Announce Type: replace-cross Abstract: Explainable AI (XAI) methods are commonly evaluated using functional correctness metrics, sometimes termed faithfulness or fidelity, which est→12 Aug 2026BREAD: Baseline-Referenced Explanations for Anomaly DiagnosisarXiv:2608.10587v1 Announce Type: new Abstract: Artificial Intelligence (AI)-based prospective anomaly detection methods are increasingly deployed in high-dimensional and nonlinear settings. Among the
CompanyDeepSeek8 recent entries6 Aug 2026MESH: Memory-Efficient Sinkhorn Optimization for Mixture-of-Experts TrainingarXiv:2608.04407v1 Announce Type: cross Abstract: Memory-efficient matrix optimizers such as Sinkhorn gradient descent remove most AdamW optimizer state for dense Transformer matrices, but direct appl→10 Aug 2026DeepSeek V4 Flash 0731 is the ‘killer app’ that is going to sell A LOT of DGX SparksHaving a ‘Killer Application’ that everyone wants to use helps sell hardware, plain and simple. DeepSeek V4 Flash 0731 isn’t an app of course, but I think it’s going to be the major catalyst for getti→11 Aug 2026I gave DeepSeek V4 Flash basic vision by training a 40M connector on 100K examplesI wanted to find out whether a huge text-only MoE could be given basic vision without retraining the language model itself. The short answer is yes. I froze DeepSeek V4 Flash and a 417M-parameter Moon→11 Aug 2026How to Ask the AI: A User Perspective Survey for Large Language Model PromptingarXiv:2608.07494v1 Announce Type: cross Abstract: AI tools like ChatGPT and DeepSeek, powered by Large Language Models (LLMs), allow users to obtain instant and effective content responses simply by t→12 Aug 2026Rethinking LLM Verification: Evidence Structure, Uncertainty, and Selective RefinementarXiv:2608.10725v1 Announce Type: new Abstract: Large language models (LLMs) often rely on shortcuts rather than systematic reasoning, raising safety concerns in medical applications. Allowing models →12 Aug 2026Measuring Semantic Abstractness of SAE Features via NonlocalityarXiv:2608.10537v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) have helped uncover mechanistic explanations for LLM behaviours such as reasoning, jailbreaking etc., via understanding the c→12 Aug 2026CohereLabs/North-Micro-Vision-Instruct · Hugging FaceNorth Micro Vision Instruct is a 2.4B-parameter open-weight vision-language model with native-resolution image support, released under the Apache 2.0 license. It is designed as a compact foundation fo→12 Aug 2026CHORUS: Complementary Experts for High-Coverage Testbench Stimulus GenerationarXiv:2608.10090v1 Announce Type: new Abstract: Large language models (LLMs) have advanced code generation, where executable feedback provides a more reliable learning signal than textual imitation al
CompanyNVIDIA8 recent entries10 Aug 2026DeepSeek V4 Flash 0731 is the ‘killer app’ that is going to sell A LOT of DGX SparksHaving a ‘Killer Application’ that everyone wants to use helps sell hardware, plain and simple. DeepSeek V4 Flash 0731 isn’t an app of course, but I think it’s going to be the major catalyst for getti→11 Aug 2026v0.32.8Muse Glimmer Muse Glimmer is now available on all platforms. Muse Glimmer can power coding agent applications such as Claude Code, Codex, Pi and more, as well as long-running personal assistants such →11 Aug 2026NVIDIA and Local AI Community Fuel Open Source Models and Intelligent AgentsThe open source ecosystem is making it easier for AI enthusiasts and developers to build, customize and run increasingly capable agents locally. Throughout August, NVIDIA is celebrating the partners a→11 Aug 2026FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal ApplicationsarXiv:2607.18171v2 Announce Type: replace Abstract: Real-time multimodal applications, including voice agents and interactive video generation, compose heterogeneous models into pipelines whose effici→11 Aug 2026ERF-GS: Reconstructing Fast Motion from Disjoint Event-RGB ViewpointsarXiv:2608.08531v1 Announce Type: new Abstract: Deep learning-driven representations such as neural radiance fields (NeRFs) and 3D Gaussian splatting (3DGS) have revolutionized the field of dynamic 3D→12 Aug 2026LiquidAI/LFM2.5-VL-3B · Hugging FaceLFM2.5-VL-3B is a multimodal variant of LFM2.5, a family of hybrid models designed for on-device deployment. It builds on LFM2-VL-3B with further mid- and post-training. LFM2.5-VL-3B can process both →12 Aug 2026How to Choose Full-Stack Observability for NVIDIA AI FactoriesA full‑stack observability framework for NVIDIA AI factories links telemetry from compute, networking, storage, orchestration and application layers using specialized tools (DCGM, NVSM, UFM, NetQ, NMX→12 Aug 2026Hand-Written PTX Tensor-Core GEMM Kernels: A Multi-Precision Study on NVIDIA L4arXiv:2608.10103v1 Announce Type: cross Abstract: High-performance Tensor Core kernels rely on a low-level PTX pipeline built from asynchronous data movement with cp.async, warp-level matrix loads wit