OpenAI just released its answer to Claude Mythos
OpenAI is launching Daybreak, an AI initiative focused on detecting and patching vulnerabilities before attackers find them. Daybreak uses the Codex Security AI agent that launched in March to create
Knowledge catalogue
OpenAI is launching Daybreak, an AI initiative focused on detecting and patching vulnerabilities before attackers find them. Daybreak uses the Codex Security AI agent that launched in March to create
arXiv:2605.07815v1 Announce Type: cross Abstract: Muon improves neural-network training by orthogonalizing matrix-valued updates, but it leaves each layer's update magnitude controlled mostly by a glo
arXiv:2602.03490v2 Announce Type: replace Abstract: Adaptive cognition requires structured internal models of objects and their relations. Predictive neural networks are often proposed to learn such w
arXiv:2605.07154v1 Announce Type: new Abstract: Referring Audio-Visual Segmentation (Ref-AVS) seeks to localize and segment target objects in video frames based on visual, auditory, and textual referr
arXiv:2605.07938v1 Announce Type: new Abstract: Single-cell representation learning (SCRL) from gene expression data offers a way to uncover the complex regulatory logic underlying cellular function.
arXiv:2605.06830v1 Announce Type: cross Abstract: Protein language models (pLMs) produce per-residue representations that capture evolutionary and structural information, yet their mean-pooled sequenc
arXiv:2605.07334v1 Announce Type: new Abstract: Video Reasoning Segmentation (VRS) aims to segment target objects in videos based on implicit instructions that convey human intent and temporal logic.
arXiv:2605.08019v1 Announce Type: new Abstract: Humans rapidly learn abstract knowledge when encountering novel environments and flexibly deploy this knowledge to guide efficient and intelligent actio
arXiv:2605.07134v1 Announce Type: cross Abstract: Web agents perceive web pages through an observation space, yet its granularity has remained an underexamined design choice. Existing work treats obse
arXiv:2602.19974v2 Announce Type: replace Abstract: Recent advancements in image generation have achieved impressive results in producing high-quality images. However, existing image generation models
arXiv:2605.07181v1 Announce Type: new Abstract: Sparse-view satellite image surface reconstruction remains highly challenging, fundamentally because the reliability of multi-view matching under satell
arXiv:2602.15189v2 Announce Type: replace-cross Abstract: Producing output that conforms to a specified JSON schema underlies tool use, structured extraction, and knowledge base construction in modern
arXiv:2605.07530v1 Announce Type: new Abstract: The Danish Technological Institute (DTI) focuses on transferring advanced technologies (including robots) to the industry and the public sector. One key
arXiv:2502.01941v3 Announce Type: replace-cross Abstract: While Key-Value (KV) cache compression is essential for efficient LLM inference, current evaluations disproportionately focus on sparse retrie
arXiv:2603.24755v2 Announce Type: replace-cross Abstract: Software development is iterative, yet agentic coding benchmarks hide design issues through their single-shot setup. Recent iterative benchmar
arXiv:2605.07161v1 Announce Type: new Abstract: AI agents are increasingly used to diagnose and mitigate failures in production systems, known as agentic Site Reliability Engineering (SRE). Current SR
arXiv:2603.21824v2 Announce Type: replace-cross Abstract: Steel surface defect analysis is critical for industrial quality control, yet existing benchmarks rely primarily on label-only annotations, li
arXiv:2605.07676v1 Announce Type: new Abstract: Standard flow matching scales well but typically relies on an unstructured source distribution, limiting its ability to learn interpretable latent struc
arXiv:2605.06729v1 Announce Type: cross Abstract: We present the EDelta-MHC-Geo Transformer, a novel architecture that unifies Manifold-Constrained Hyper-Connections (mHC), Deep Delta Learning (DDL),
arXiv:2605.08060v1 Announce Type: cross Abstract: Context window expansion is often treated as a straightforward capability upgrade for LLMs, but we find it systematically fails in multi-agent social
arXiv:2605.07158v1 Announce Type: cross Abstract: Vector search and retrieval-augmented generation (RAG) rest on the assumption that cosine similarity between text embeddings reflects conceptual relat
arXiv:2605.07935v1 Announce Type: new Abstract: We present TraceFix, a verification-first pipeline for Large Language Model (LLM) multi-agent coordination. An agent synthesizes a protocol topology as
arXiv:2605.07772v1 Announce Type: new Abstract: Transformers perform inference by iteratively transforming token representations across layers. This layerwise computation has been studied empirically,
TIL: Using LLM in the shebang line of a script Kim_Bruning on Hacker News: But seriously, you can put a shebang on an english text file now (if you're sufficiently brave) [...] This inspired me to loo
arXiv:2602.16571v2 Announce Type: replace Abstract: Large-scale sharing of dialogue data is key to advancing the science of teaching and learning, yet rigorous de-identification remains a major barrie
arXiv:2605.06708v1 Announce Type: cross Abstract: Visual text compression (VTC) promises efficient long-context processing by rendering text into an image and re-encoding it with a vision-language mod
📣We're calling for ambassadors! Whether you're a developer with great technical taste or a local community leader who loves bringing people together, we'd love to have you join us. Visit the website b
arXiv:2605.07756v1 Announce Type: cross Abstract: Modern deep models are often pretrained on large-scale data with missing labels using composite objectives, where the relative weights of multiple los
arXiv:2605.06731v1 Announce Type: cross Abstract: Personalized LLM agents maintain persistent cross-session state to support long-horizon collaboration. Yet, this persistence introduces a subtle but c
arXiv:2605.05405v2 Announce Type: replace Abstract: Semantic search of Earth observation archives remains challenging. Visual foundation models such as CLAY produce rich embeddings of satellite imager
Sam Altman comments on a statement by @roydanroy, suggesting that the technological capabilities or developments being discussed would not have been possible prior to GPT-5.5. The post implies that GP
Downloading now... 1M token context window with supposedly usable coding agent capability all on a 128GB Macbook Pro is 🤯 🚨 OPEN SOURCE AI IS LITERALLY UNSTOPPABLE 🚨 The legendary founder of Redis (An
Hot take on METR’s new graph that so many people are flipping about today. • Claude Code is a real advance; Mythos probably builds on some of what is learned there. But… • If you read the graph carefu
This is confused, but popular. Popular because it tells a bunch of people what they want to hear. Confused for a couple reasons: first, Mythos probably isn’t a pure LLM. (Claude Code isn’t, and it pro
Yann LeCun closed 1.03B for AMI Labs on March 10. Three days later, this paper dropped from his NYU collaborators. 15M parameters. Single GPU. A few hours of training. LeWorldModel is the first JEPA t
Chain of thought monitors are a key layer of defense against AI agent misalignment. To preserve monitorability, we avoid penalizing misaligned reasoning during RL. We found a limited amount of acciden
This post describes a project where the developer created an autonomous AI agent capable of modifying and executing its own source code across a 7-day development period, integrating multiple language
we'd like to help companies secure themselves and we think it's important to start work on this quickly Today, we're rolling out GPT‑5.5‑Cyber in limited preview to defenders responsible for securing
arXiv:2603.09789v2 Announce Type: replace Abstract: Accurate financial volatility forecasting is crucial but challenged by the non-linear, highly correlated nature of market data. Recently, quantum co
arXiv:2605.04055v1 Announce Type: new Abstract: Adaptive optimizers like AdamW apply uniform hyperparameters across all parameter groups, ignoring heterogeneous optimization dynamics across layers and
arXiv:2605.05079v1 Announce Type: new Abstract: Video sequence capturing through refractive dynamic media, such as a turbulent air or water surface, often suffer from severe geometric distortions and
arXiv:2605.04066v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) is an essential paradigm that enhances the reasoning capabilities of Large Language Models (LLMs).
arXiv:2507.23501v2 Announce Type: replace Abstract: Ensembles are ubiquitous in off-policy actor-critic learning, yet their efficacy depends critically on how they are aggregated. Current methods typi
arXiv:2404.06230v3 Announce Type: replace Abstract: In federated learning (FL), profiling and verifying each client is inherently difficult, which introduces a significant security vulnerability: mali
arXiv:2507.12768v2 Announce Type: replace Abstract: Learning generalizable manipulation policies hinges on data, yet robot manipulation data is scarce and often entangled with specific embodiments, ma
arXiv:2605.03339v1 Announce Type: new Abstract: Solving large-scale CVRP (LSCVRP) with hundreds to thousands of nodes remains difficult for even state-of-the-art solvers. Divide-and-conquer can scale
arXiv:2605.04299v1 Announce Type: new Abstract: Scene understanding is a vital part of autonomous driving systems, which requires the use of deep learning models. Deep learning methods are intrinsical
arXiv:2505.00020v2 Announce Type: replace Abstract: Using a legally obtained dataset of 34 copyrighted O'Reilly Media books, we apply the DE-COP membership inference attack method to investigate wheth
arXiv:2605.05014v1 Announce Type: new Abstract: Autonomous driving must operate across diverse surfaces to enable safe mobility. However, most driving datasets are captured on well-paved flat roads. M
arXiv:2605.04980v1 Announce Type: cross Abstract: Activation-based steering provides control of LLM behavior at inference time, but the dominant paradigm reduces each concept to a single direction who
arXiv:2605.04073v1 Announce Type: new Abstract: Bail decisions present a fundamental challenge for data-driven decision support systems. When bail is denied, the counterfactual outcome of whether the
arXiv:2605.04405v1 Announce Type: new Abstract: Driven by the rapid development of generative AI models, deepfake detectors are compelled to undergo periodic recalibration to capture newly developed s
arXiv:2605.03227v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated strong capabilities in natural language understanding and reasoning. However, their ability to perform ex
arXiv:2605.05034v1 Announce Type: new Abstract: Despite the strong performance of Convolutional Neural Networks (CNNs) in disease classification, their effectiveness often depends on access to large a
arXiv:2604.01496v2 Announce Type: replace-cross Abstract: We introduce SWE-ZERO to SWE-HERO, a two-stage SFT recipe that achieves state-of-the-art results on SWE-bench by distilling open-weight fronti
arXiv:2605.04535v1 Announce Type: new Abstract: Inferring continuum models directly from video is hampered by two facts: the recorded field is uncalibrated image intensity rather than a physical state
arXiv:2605.04770v1 Announce Type: new Abstract: While zero-shot appearance-based 3D gaze estimation offers significant cost-efficiency by directly mapping RGB images to gaze vectors, its reliability i
arXiv:2605.04449v1 Announce Type: new Abstract: Dialogue State Tracking (DST) requires precise extraction of structured information from multi-domain conversations, a task where Large Language Models
arXiv:2603.26178v2 Announce Type: replace Abstract: We introduce the Geometric Evolution Graph Convolutional Network (GEGCN), a novel framework that enhances graph representation learning through expl
arXiv:2605.03410v2 Announce Type: new Abstract: Out-of-distribution (OOD) detection identifies test samples that fall outside a model's training distribution, a capability critical for safe deployment