Shieldstral
arXiv:2607.25857v1 Announce Type: new Abstract: We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7imes its size on text s
Knowledge catalogue
arXiv:2607.25857v1 Announce Type: new Abstract: We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7imes its size on text s
arXiv:2607.25063v1 Announce Type: new Abstract: Developers judge a model checkpoint by how it behaves. After supervised fine-tuning (SFT), two checkpoints that perform about the same across relevant b
arXiv:2607.25026v1 Announce Type: cross Abstract: In this work, we present a simulation-based parameter estimation framework for a model defined by a computational simulation of a physical system. We
Enterprise data integration and automation firm SnapLogic Inc. today announced a significant update to SnapGPT, the company’s artificial intelligence copilot for enterprise data automation, transformi
arXiv:2607.25219v1 Announce Type: new Abstract: Social navigation has progressed from simplified 2D environments toward a more general vision-based setting, in which a robot needs to achieve socially
Bloomberg: Source: the European Commission plans to designate OpenAI's ChatGPT and Roblox as “very large online platforms” under the DSA as soon as August — OpenAI's ChatGPT and video game company Rob
arXiv:2607.24802v1 Announce Type: cross Abstract: This paper presents our system for Task 3 of the CLEF 2026 CheckThat! Lab, which focuses on generating full fact-checking articles from claims, veraci
Madhumita Murgia / Financial Times: Sources: Google DeepMind has reassigned the majority of the original authors of the AlphaFold papers; about a quarter of the papers' full-time authors have left — L
arXiv:2607.25413v1 Announce Type: new Abstract: Transaction propensity prediction in B2B e commerce presents unique challenges distinct from B2C contexts, primarily due to the heterogeneous procuremen
arXiv:2607.24787v1 Announce Type: new Abstract: Sparse Mixture-of-Experts (MoE) models expand foundation model capacity through conditional expert activation, but their full expert pools remain diffic
arXiv:2607.25790v1 Announce Type: cross Abstract: This paper proposes a novel physics-guided spectral deep operator network, termed SpectONet, for solving Euler-Bernoulli beam (EBB) vibration problems
arXiv:2607.25074v1 Announce Type: cross Abstract: Synthetic control (SC) matches a treated unit's pre-treatment trajectory to a weighted combination of donor units. We study Spectral SC, which instead
arXiv:2607.25333v1 Announce Type: cross Abstract: Specula is a push-button agentic system that generates high-quality formal specifications for large, complex system code and uses the specifications f
arXiv:2607.25816v1 Announce Type: new Abstract: Large language model agents often spend substantial wall-clock time waiting for tool call results. Tool-call speculation can hide this latency by predic
arXiv:2607.25716v1 Announce Type: new Abstract: Federated learning (FL) enables privacy-preserving training of automatic speech recognition (ASR) systems across distributed data sources, yet its appli
arXiv:2607.26052v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) variants of Low-Rank Adaptation (LoRA) route every token to a fixed number of experts k. Tokens differ in how uncertain the mod
arXiv:2607.24953v1 Announce Type: cross Abstract: Reducing training precision is a key lever for improving the e ciency of large language model (LLM) training, but pushing beyond FP8 to 4-bit oating p
arXiv:2602.20555v2 Announce Type: replace-cross Abstract: The tremendous success of Transformer models in fields such as large language models and computer vision necessitates a rigorous theoretical i
arXiv:2607.24777v1 Announce Type: new Abstract: Architected metamaterials derive their functions from structure, creating vast opportunities to program physical responses through topology design. Howe
arXiv:2607.24863v1 Announce Type: new Abstract: Control methods for emerging vehicle chassis architectures are important for autonomous driving near handling limits. Unlike conventional drift control,
arXiv:2607.25880v1 Announce Type: cross Abstract: LLM provenance testing asks whether a suspect LLM belongs to the same lineage as a source. Existing black-box methods largely infer this relationship
arXiv:2607.24827v1 Announce Type: cross Abstract: We prove two lower bounds for the first order oracle complexity of minimizing a d-dimensional 1-Lipschitz convex function over the unit ball with m bi
arXiv:2607.25268v1 Announce Type: cross Abstract: Ranking is a fundamental component of modern information access systems. Reinforcement learning (RL) provides a flexible framework for directly optimi
Super interesting new work from NVIDIA. (bookmark it) They suggest building agents as Python objects. Very cool idea and I think it could a lot with agent reliability. More below: Agent development to
arXiv:2511.16618v2 Announce Type: replace Abstract: Surgical scene understanding demands temporally consistent tracking of instruments and tissues. For clinical use, such tracking should generalize to
Israeli runtime security startup Sweet Security Ltd. today unveiled Agentic AI Blocking, a capability that stops artificial intelligence agents mid-action in production when they do something they wer
arXiv:2607.25182v1 Announce Type: cross Abstract: The ability to retrieve relevant tables for answering questions is a key task for structured information retrieval. Multi-stage retrieval systems rely
arXiv:2507.10643v4 Announce Type: replace-cross Abstract: Post-hoc model-agnostic local attribution (LA) methods have been widely adopted to explain opaque AI models by quantifying feature-wise contri
arXiv:2607.25337v1 Announce Type: new Abstract: Joint-Embedding Predictive Architectures (JEPAs) learn world models by predicting in representation space rather than reconstructing pixels, making them
Today’s data lakehouse is no longer mere data repository, but increasingly a system of action, actively executing tasks via always-on, autonomous AI agents. Rather than waiting for static reports, the
arXiv:2607.25425v1 Announce Type: new Abstract: Capture the Flag (CTF) competitions are among cybersecurity's most effective training grounds, developing practical skill across cryptography, web explo
arXiv:2607.24767v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) systems have emerged as a powerful process for allowing large language models (LLMs) to retrieve relevant informa
Capitalizing on agentic AI depends on the right infrastructure investments, yet enterprises struggle to reduce cloud costs. Vultr VX1™ Cloud Compute presents a way to free up budget for CPUs and GPUs
On the CPU, batch 1 is memory bandwidth bound. But if token/s = bandwidth / (bytes_per_weight * active_weights_per_token) the total number of parameters doesnt slow down the generation speed. So build
arXiv:2607.25570v1 Announce Type: cross Abstract: The development of autonomous vehicles (AVs) usually relies heavily on data-driven artificial intelligence (AI) models that require large volumes of s
The median task consumed nearly 6x more tokens in Claude Code than in Kimi Code: - 61k in Kimi Code - 67k in Hermes - 340k in Claude Code At K3's 3/M input rate (input tokens make up roughly 95% of ag
The number of times I have had to tell Opus 5 and Fable 5 to talk to me like an excited teenager and not a computer science phd is high. Was going through a Google auth thing and got multiple paragrap
The openai agent sandbox escape actually has very real implications for the diffusion of AI in the enterprise. The incident showed the power and capability of agents, and the need to harden systems an
arXiv:2604.14990v2 Announce Type: replace Abstract: The prospect of Artificial General Intelligence (AGI) is increasingly driving institutional decisions, and alignment of AGI is a hard problem. The c
The problem: agent workflows alternate between GPU-heavy reasoning and GPU-idle waiting on tools. Run hundreds concurrently and their KV caches fight for memory. Engines evict on a dumb LRU policy, ev
The root cause: request-level engines never see that a series of LLM calls belongs to one longer workflow. ThunderAgent adds that missing view. It treats each agent workflow as a schedulable program,
arXiv:2605.04330v2 Announce Type: replace Abstract: We investigate the scaling properties of implicit deductive reasoning over Horn clauses in depth-bounded Transformers. By systematically decorrelati
Quick clarification before opening the bag: every image here is fictional and AI-generated. These are not real archival discoveries or claims about hidden history. This series grew out of two earlier
arXiv:2607.25253v1 Announce Type: new Abstract: Online recommendation has traditionally taken place after a user enters a platform, which determines the candidate pool and the ranking shown to the use
OpenAI announced the release of GPT‑5.6 Sol after deployment, incorporating optimizations across its stack to enhance run‑time efficiency. The update delivers a roughly 20 % reduction in serving costs
This is a willfully misleading narrative from OpenAI. Sam’s earnest expressions are being deployed, again, to misdirect. If there’s a need to “pace AI development”, the OAI hack incident isn’t any evi
arXiv:2607.24781v1 Announce Type: cross Abstract: RAG systems rely on chunking, which destroys structural information in documents. Existing heading-based retrieval (Jeong et al., 2025) requires multi
Thrilled to have @gabepereyra speak at our @sequoia event tmrw on OWN YOUR AI: how to build your own Lab as an application company. Also featuring @FireworksAI_HQ @mercor_ai @LangChain @trajectorylabs
ThunderAgent is a program-aware scheduler for agentic inference. By treating each agent workflow as a schedulable program, it eliminates KV cache thrashing to deliver more than 2x single-node throughp
arXiv:2607.25894v1 Announce Type: new Abstract: Recent diffusion models have achieved remarkable realism in facial image synthesis, posing growing challenges to artificial intelligence-generated conte
arXiv:2607.17761v2 Announce Type: replace-cross Abstract: Recently, speech deepfake detection (SDD) has achieved significant progress. However, its robustness evaluation remains largely confined to co
arXiv:2607.24750v1 Announce Type: new Abstract: Large Language Models (LLMs) are temporally overexposed: trained on vast contemporary corpora, they encode present-day concepts that make them unreliabl
Today we’re open-sourcing Numbat, an agent-detection and response layer that is designed to work across agent harnesses. Numbat gives security teams visibility into agent activity, with controls to bl
arXiv:2607.10016v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used as backbone architectures for recommender systems because of their strong sequence modeling
arXiv:2607.24865v1 Announce Type: cross Abstract: Large-scale recommendation systems face 'Memory Wall' bottlenecks due to massive, dense embedding tables. While generative retrieval uses discrete tok
arXiv:2607.25718v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly rely on invoking external tools to complete real-world tasks. Tool retrieval, which selects a small tas
arXiv:2607.25216v1 Announce Type: cross Abstract: Semantic ID-based generative recommendation tokenizes each item into a sequence of discrete semantic IDs and predicts the next item by generating sema
arXiv:2607.25305v1 Announce Type: new Abstract: Macroareas are geographical areas used in typological research for grouping variables of interest. In linguistic typology, languages in a given macroare
arXiv:2607.25446v1 Announce Type: new Abstract: Multi-agent frameworks built on large language models (LLMs) routinely entangle three logically distinct concerns: who is on the team (organization), ho
arXiv:2607.25914v1 Announce Type: new Abstract: Autonomous Network Levels 4-5 require AI agents to invoke tools across vendor boundaries without human oversight, yet existing management standards lack