Subliminal Learning is Non-Semantic Distillation
arXiv:2608.05734v1 Announce Type: new Abstract: Subliminal Learning (SL) is a surprising type of generalization displayed by modern language models. It allows the transfer of a bias or behavior from a
Knowledge catalogue
arXiv:2608.05734v1 Announce Type: new Abstract: Subliminal Learning (SL) is a surprising type of generalization displayed by modern language models. It allows the transfer of a bias or behavior from a
arXiv:2608.05785v1 Announce Type: cross Abstract: Multilingual text embedding models are commonly adapted using a single training objective across diverse tasks, despite different tasks requiring fund
arXiv:2608.05699v1 Announce Type: new Abstract: Humans understand anomalous events through a coherent perceptual process in which they identify the focal instance, follow its behavior as the event unf
arXiv:2608.05389v1 Announce Type: new Abstract: Background: Accurate glioma subregion delineation is important for radiotherapy planning and longitudinal monitoring, but manual contour correction is t
Alibaba’s Qwen team announced the new Qwen‑3.8‑Max model and thanked users for extensive testing, noting its ability to provide highly detailed world observations. A community member highlighted that
arXiv:2608.05982v1 Announce Type: new Abstract: Inadequate target--disease linkage accounts for 40--50% of Phase~II efficacy failures, so anticipating which programmes will advance would let sponsors
The backgrounds were generated with MiniMax M3. Everything else, including the full game design, was done with DeepSeek Flash. Most importantly, all of this was done inside Hermes Agent. You won't bel
arXiv:2608.06370v1 Announce Type: new Abstract: Tool use transforms LLMs into agents that act beyond their training data, and for code-capable models, programmatic tool calling extends this further by
The cloud was always a cat. Qwen3.8-Max just saw it first. 😼☁️ Try it yourself! Qwen 3.8 Max is actually impressive Sent a sky pic to it along with Claude Opus 5, Kimi K3, and GPT 5.6 Sol, and asked t
arXiv:2608.05160v1 Announce Type: new Abstract: We introduce the Ignition Index (I), a validated scalar metric that operationalizes Global Workspace Theory's (GWT) all-or-none ignition prediction in t
arXiv:2608.06361v1 Announce Type: new Abstract: Real-world video benchmarks provide broad coverage, but their fixed clips entangle event count, rate, duration, and visual complexity, making failure mo
arXiv:2608.06065v1 Announce Type: new Abstract: GUI agents are commonly trained offline from successful interaction trajectories. Standard training decomposes each trajectory into prefix-action pairs:
this talk by openai researchers going through hugging face incident is totally insane, so much to unpack openai only realized it was their agent who hacked hugging face infra while asking hf to revoke
arXiv:2608.06177v1 Announce Type: new Abstract: Binary neural networks are very attractive for constrained deployment, enabling small footprint and low-power inference. For binary activations, the dot
arXiv:2608.06252v1 Announce Type: cross Abstract: Deaf and hard-of-hearing people in Bangladesh communicate mainly through Bangla Sign Language (BdSL). Automatic BdSL recognition on personal devices c
arXiv:2608.06346v1 Announce Type: new Abstract: LLM-based agentic systems have shown remarkable capabilities in complex domains, while suffering from cascading errors and difficulty in debugging. Crit
arXiv:2501.02140v2 Announce Type: replace-cross Abstract: This paper introduces Tree-NET, a novel framework for medical image segmentation that leverages bottleneck supervision to enhance both segment
arXiv:2608.06223v1 Announce Type: new Abstract: While deep learning models, particularly transformer-based architectures, have shown impressive performance in time series forecasting, the application
turns out you can get indirect prompt injection to ~0 on unseen attacks if you stack enough layers (model training + input probes + a classifier checking intent). didn't expect that a year ago. auto m
arXiv:2608.06331v1 Announce Type: cross Abstract: From natural-language query interfaces to automated report generation, data analysis tools need a description of the data: the real-world entities it
arXiv:2608.05729v1 Announce Type: new Abstract: As capabilities rapidly increase, AI agents can move from running inside one app to acting across a user's devices over time. Yet existing agent systems
Modern enterprises possess a vast amount of unstructured data, yet they frequently encounter significant challenges in managing and extracting value from it. Historically, unlocking the insights hidde
arXiv:2608.05153v1 Announce Type: cross Abstract: GraphRAG underperforms vector RAG on citation precision in many reports, but where and why have remained corpus-bound. We present a triple-robustness
upgraded my stack, and i can now work on almost anything from anywhere hands free: - talk to chief of staff (via remote codex voice or text) - chief assigns tasks to managers of various projects - man
arXiv:2608.05545v1 Announce Type: cross Abstract: Generative AI used as a capable servant has greatly accelerated intellectual work, but it also risks eroding human epistemic agency by encouraging unc
arXiv:2608.05485v1 Announce Type: new Abstract: Evaluating generated videos remains challenging because existing benchmarks rely on fixed evaluation content, cover only a subset of generation and edit
arXiv:2504.13621v2 Announce Type: replace Abstract: Visual grounding associates textual descriptions with objects in an image. Conventional methods target third-person image inputs and named object qu
arXiv:2608.05776v1 Announce Type: new Abstract: Autoregressive continuation provides a natural path toward minute-scale audio-visual generation by repeatedly extending a short-window generator conditi
We analyzed DeepSeek-V4 Flash-0731 vs. GPT-5.6 Luna on software engineering tasks using DeepSWE. DeepSeek-V4 Flash-0731 delivers 80% of Luna’s performance at roughly 1/6 the cost. More insights in the
arXiv:2608.06202v1 Announce Type: cross Abstract: Large language model (LLM) benchmark evaluations are routinely used to support claims about model safety, reliability, and deployment readiness. Yet m
arXiv:2606.14299v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) such as CLIP have become a standard backbone for open-vocabulary recognition, yet their zero-shot predictions remain v
arXiv:2608.06057v1 Announce Type: new Abstract: Tool-calling agents infer task state from accumulated dialogue and tool traces. In persistent interactions, however, historical traces may remain struct
arXiv:2608.05810v1 Announce Type: new Abstract: Self-evolving agents accumulate capability by distilling reusable skills from their execution trajectories, but we find this process is not monotonic: p
arXiv:2608.05163v1 Announce Type: new Abstract: A common assumption holds that switching to a non-English language makes a multilingual RAG system easier to attack for personal information. We test th
arXiv:2606.21155v2 Announce Type: replace Abstract: Attorneys, judges, and pro se filers increasingly use AI to draft legal documents, yet these tools frequently fabricate citations. Despite predictio
You can play it here: https://simonw.github.io/raccoon-heist-codex/ For comparison, here's Fable 5 + Claude Code's game, built from the exact same prompt https://x.com/simonw/status/208508951822360205
arXiv:2506.03259v3 Announce Type: replace Abstract: Purpose: To compare five lightweight open-weight large language models (LLMs) with a rule-based algorithm (RBA) and fine-tuned RadBERT for zero-shot
Following up on yesterday's post about running everyone's faves on 2 x 16gb cards while maximizing performance and KV. Previous post data used abandoned Cu130 VLLM image. Stats here are done on cu129-
5.6 Sol much better in chat now and unlimited text chat for free users! We’re making better intelligence easier to access in ChatGPT for everyone: - GPT-5.6 Sol now powers both Instant and deep reason
arXiv:2608.04451v1 Announce Type: cross Abstract: We give a negative solution to MAIS-O60. We first construct an example in which an initially active ReLU neuron becomes completely inactive in finite
arXiv:2602.06052v4 Announce Type: replace-cross Abstract: Research in artificial intelligence is shifting from model innovations and benchmark scores towards problem definition and rigorous real-world
arXiv:2608.04215v1 Announce Type: cross Abstract: The growing diversity of code clone types, from syntactic copies to cross-language semantic clones to AI-generated duplicates, has created a fragmenta
arXiv:2608.04792v1 Announce Type: new Abstract: Accurate estimation of Above-Ground Biomass (AGB) from satellite imagery is essential for the large-scale monitoring of carbon stocks, yet it remains a
arXiv:2608.04581v1 Announce Type: new Abstract: Recent advances in 4D Gaussian Splatting (4DGS) enable high-fidelity, real-time spatiotemporal rendering, but expose a fundamental trade-off between mot
arXiv:2608.04682v1 Announce Type: cross Abstract: Coding agents powered by large language models (LLMs) are increasingly adopted in software engineering (SWE) scenarios, capable of fixing a specific b
arXiv:2608.04305v1 Announce Type: new Abstract: Risk-aware Q-learning (RaQL) provides a model-free, two-timescale estimator for dynamic risk objectives, but its finite-budget behavior remains fragile:
The intersection of medicine and AI has led to remarkable innovations. However, developers now face the thorny challenge of building robust medical AI tools that have been tested and evaluated on dive
arXiv:2608.04061v1 Announce Type: new Abstract: Utility poles are an essential part of the infrastructure used to support power distribution systems and other critical public services. Their regular i
arXiv:2608.04314v1 Announce Type: cross Abstract: Once visual content enters an AI pipeline, its owner often retains little technical control over how it is used. Legal and regulatory remedies can add
arXiv:2608.04190v1 Announce Type: new Abstract: Deploying pre-trained perception models in novel environments degrades their accuracy under distributional shift, and assembling them alone does not rec
In the modern data landscape, query performance tuning and managing system price-performance is challenging, especially as the number of agentic workloads increase. Even for experienced developers and
arXiv:2608.04917v1 Announce Type: new Abstract: Biological visual systems can perceive depth from monocular vision flow, continuously integrating temporal visual cues while maintaining a balance betwe
An AI model from Meta also hacked another company during testing Stop me if you've heard this one before: An AI model from the parent company of Facebook and Instagram hacked into another company’s sy
A small update to Sir Shortoken. Sir Shortoken already had Quick, Balanced, Deep, Bullets, and Aggressive Bullets. I wanted something between Bullets and normal prose. So I added LELP-S+ (Less English
Anthropic is hiring a custom silicon team to design proprietary chips that will power its Claude models, while still planning a multi‑chip strategy that mixes internally designed hardware with compone
anyone confused about neurosymbolic AI—and its recent enormous victory—should read this. complete and total vindication for what I have been saying here all along. (see also my essays on Claude Code a
The tweet is a question from user Simon Willison (posted on 6 Aug 2026) asking which OpenAI API model corresponds to the ChatGPT “GPT‑5.6 Instant” version. No answer or clarification is included in th
arXiv:2608.05144v1 Announce Type: new Abstract: Long-horizon reasoning requires an agentic runtime that can persist when evidence supports its current approach and pivot when measurements reveal failu
arXiv:2607.16060v2 Announce Type: replace Abstract: Artistic charts combine data visualization with expressive marks, textures, and typography, but they are difficult for image generators: an output i
arXiv:2608.04213v1 Announce Type: new Abstract: Existing studies on self-supervised learning for white-box networks typically decouple the derivation of white-box networks via optimization algorithms