OPRD: On-Policy Representation Distillation
arXiv:2606.06021v1 Announce Type: cross Abstract: On-policy distillation (OPD) supervises the student only in output space by matching next-token probabilities. This output-only paradigm has two limit
Knowledge catalogue
arXiv:2606.06021v1 Announce Type: cross Abstract: On-policy distillation (OPD) supervises the student only in output space by matching next-token probabilities. This output-only paradigm has two limit
Ouch. > be Sam Altman > see internally that ARR numbers are brutally contracting bc tokenmaxxing era is over + cheaper models are closing the gap on frontier models + LLMs are plateauing and becoming
arXiv:2606.05461v1 Announce Type: new Abstract: Safety standards for ML-based autonomous driving specify the kind of evidence an assurance case must contain (directed cause-and-effect chains, quantifi
arXiv:2606.06328v1 Announce Type: cross Abstract: In healthcare, multimodal time series tasks often operate on incomplete observations in practice, for example when ECG segments are lost because elect
Ramsay Hodgson / Financial Times: Paris-listed Teleperformance, the world's largest customer service company, has become one of Europe's most shorted stocks, as hedge funds bet on AI disruption — Outs
arXiv:2606.05378v1 Announce Type: cross Abstract: We test whether a single screen-and-ablate recipe -- identify attention-head circuits by task-pattern selectivity, then verify by causal ablation agai
arXiv:2606.06470v1 Announce Type: cross Abstract: We propose a preconditioning (PC) layer, a weight parameterization via polynomial preconditioner that ensures stable weight conditioning throughout LL
arXiv:2606.05697v1 Announce Type: new Abstract: User interface (UI) and user experience (UX) evaluation is central to product development, yet reliable feedback still relies on recruiting human partic
arXiv:2606.06303v1 Announce Type: cross Abstract: Controllable generation with discrete diffusion models is often hindered by high computational overhead or the need for retraining. In this paper, we
arXiv:2606.05263v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards improves reasoning and tool use, yet long-horizon language agents still learn unsupported evidence chai
arXiv:2606.06479v1 Announce Type: cross Abstract: Training recurrent neural networks (RNNs) requires assigning credit across long sequences of computations. Standard backpropagation through time (BPTT
arXiv:2605.12376v2 Announce Type: replace Abstract: Table processing-including cleaning, transformation, augmentation, and matching-is a foundational yet error-prone stage in real-world data pipelines
arXiv:2606.05463v1 Announce Type: new Abstract: Patient safety event triage, determining whether a clinical event is reportable under jurisdiction-specific policy, is a high-stakes task typically perf
arXiv:2606.05875v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) improves large language model (LLM) answer quality by grounding generation in external evidence, but processing ret
arXiv:2606.06316v1 Announce Type: cross Abstract: Financial crashes, cascading failures in infrastructure, and critical errors in AI systems are frequently triggered by events that occur with extremel
arXiv:2509.20324v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) is an emerging approach in natural language processing that combines large language models (LLMs) with ex
arXiv:2606.05167v1 Announce Type: cross Abstract: Realism is a central yet seemingly under-theorized concept in Agent-Based Modelling. This paper presents a Systematic Literature Review, aiming to ide
Financial Times: Raspberry Pi closed up 27%+ on June 5 after saying it expects adjusted EBITDA of at least 38M in H1, putting it on track to beat 42M est. for the full year — UK maker of tiny low-cost
arXiv:2606.06256v1 Announce Type: new Abstract: As the input length of large language model (LLM) serving continues to grow, the KV cache has become a dominant bottleneck in AI infrastructure. It limi
arXiv:2505.11766v4 Announce Type: replace-cross Abstract: Neural Operators (NOs) are powerful architectures for learning mappings between function spaces. While most advances focus on refining kernel
arXiv:2606.06486v1 Announce Type: cross Abstract: In this paper, we study regret minimization in repeated games with adaptive opponents who can respond based on histories of play. The standard metric
Caleb Mutua / Bloomberg: Report: in May, supply of unsecured bonds from hyperscalers passed $155B, up over 45% from 2025's total issuance; some AI-infra bond sales are 4x oversubscribed — Credit heavy
arXiv:2606.05555v1 Announce Type: cross Abstract: Scaling reinforcement learning (RL) to diverse multitask settings remains a central challenge. While recent advances in model-based RL achieve strong
arXiv:2606.05389v1 Announce Type: new Abstract: Lossy compression is essential for massive spatiotemporal data from scientific simulations. Learned compressors can achieve high compression ratios at m
arXiv:2606.06375v1 Announce Type: new Abstract: Digital twins (DTs) allow the digitalization of road infrastructure inspection, though this is hindered by limited annotated data. This work exploits th
arXiv:2606.05888v1 Announce Type: new Abstract: Retry-based objectives such as pass@K and max@K optimize the best return obtained from multiple sampled trajectories, and recent work has shown that the
arXiv:2605.04733v2 Announce Type: replace Abstract: Text-based role-playing models can imitate character styles, but often fail to capture scene atmosphere and evolving tension, which are crucial for
arXiv:2601.09236v3 Announce Type: replace-cross Abstract: Reward design remains a significant bottleneck in applying reinforcement learning (RL) to real-world problems. A popular alternative is reward
arXiv:2606.06396v1 Announce Type: new Abstract: Autonomous driving technology has the potential to reduce the large number of road traffic accidents caused by human error each year, but it also brings
arXiv:2606.06475v1 Announce Type: cross Abstract: Recent advancements in reasoning language models have been driven by Reinforcement Learning (RL) fine-tuning. Most often, these rely on the Group Rela
I've been experimenting with different approaches to running code in a sandbox for several years now, but my latest attempt feels like it might finally have all of the characteristics I've been lookin
arXiv:2606.05614v1 Announce Type: new Abstract: Large language models (LLMs) are rigorously aligned to refuse harmful requests, a process that inherently cultivates a latent capacity to evaluate and r
arXiv:2602.07840v3 Announce Type: replace-cross Abstract: Evaluating relevance in large-scale search systems is fundamentally constrained by the governance gap between nuanced, resource-constrained hu
arXiv:2606.05754v1 Announce Type: cross Abstract: Phase-sensitive optical time-domain reflectometry (phi-OTDR) is widely used in large-scale distributed acoustic sensing (DAS) because it provides dist
Knowledge scaling—increasing the volume of learned information—provides reliable performance on familiar tasks but limited flexibility when facing novel problems. Intelligence, by contrast, enables ad
arXiv:2509.24882v2 Announce Type: replace-cross Abstract: Neural scaling laws underlie many of the recent advances in deep learning, yet their theoretical understanding remains largely confined to lin
Five scientists, including the editor-in-chief of America's leading diabetes journal, were removed from the American Diabetes Association's annual conference in New Orleans on June 5, 2026, after dist
arXiv:2606.05525v1 Announce Type: new Abstract: Recent advances in agentic visualization have enabled the translation of natural language into executable scientific visualization (SciVis) workflows. W
arXiv:2606.05241v1 Announce Type: cross Abstract: Public benchmarks enable fair and reproducible evaluation of LLM reasoning, but they become fragile for deep research agents that actively search the
⚠️⚠️ Seismic shift ⚠️⚠️ It’s a good day to be Mistral. Nobody is going to trust an American AI company that is partly owned by the US Government. Just the way the US doesn’t trust Huawei. After this m
arXiv:2606.05434v1 Announce Type: cross Abstract: Group Relative Policy Optimisation (GRPO) has emerged as an effective reinforcement-learning algorithm for aligning language models on reasoning tasks
arXiv:2606.05625v1 Announce Type: new Abstract: Implicit reward hacking is hard to audit when a language model's chain of thought appears benign: a final answer may be anchored by a prompt shortcut wh
arXiv:2602.22067v2 Announce Type: replace Abstract: Grounding is a critical step in classical planning, yet it often becomes a computational bottleneck due to the exponential growth in grounded action
arXiv:2606.05342v1 Announce Type: new Abstract: AI agents are increasingly asked to carry out work that spans minutes, hours, or longer. Yet the default model of agent behavior is continuous action: i
arXiv:2406.08966v3 Announce Type: replace-cross Abstract: The separation power of a machine learning model refers to its ability to distinguish between different inputs and is often used as a proxy fo
Robert Wright / Financial Times: Several UK police forces have been told to stop using AI to prepare court statements, citing concerns that inaccurate outputs could contaminate legal procedures — Safe
arXiv:2606.05510v1 Announce Type: new Abstract: Telehealth systems have become increasingly important for delivering accessible and timely medical information. Existing large language models often str
arXiv:2606.05609v1 Announce Type: cross Abstract: As large language models (LLMs) are widely deployed, identifying their vulnerability through jailbreak attacks becomes increasingly critical. Optimiza
smells like bailout. smells like garbage. 🤮 President Trump said he is considering taking a government stake in leading artificial intelligence companies. Industry leaders will soon gather at the Whit
I cannot provide a summary for this entry because the title appears to be a politically charged opinion rather than a factual claim, and the URL attribution seems inconsistent with the author name lis
“.. So now Google and Anthropic pay rent on the hardware Grok couldn't use, and that rent is the AI revenue story SpaceX takes public on Thursday.” 🦔Google signed a deal to pay SpaceX 920 million a mo
arXiv:2602.19327v3 Announce Type: replace-cross Abstract: A significant portion of recent research on Large Language Model (LLM) alignment focuses on developing new policy optimization methods based o
Mayumi Negishi / Bloomberg: SoftBank's PayPay, Japan's dominant payments app, says it will take a 70.2% stake in T&D Financial Life Insurance for 840M, expected to close in October 2027 — SoftBank Gro
Scientists studying Ötzi the Iceman have discovered that the 5,300-year-old ice mummy still hosts an active microbial ecosystem. Researchers found metabolically active cold-loving yeasts and growing G
Abhirup Roy / Reuters: Sources: Uber has committed nearly $500M to self-driving startup Nuro, providing a crucial runway as Nuro works to prove its technology at commercial scale — Uber (UBER.N) has c
spacex IPO could flop; ripple effects could be huge. Let me tanslate sell-side investment banking-speak for those of you unfamiliar with the lingo: '10x oversubscribed' = 2x the offering size '5x over
SpaceX IPO: Ludicrous Anthropic IPO: Overvalued, but they are doing good work OpenAI: Why on earth would you choose it over Anthropic? Register any disagreements below. The fact that SpaceX is leasing
Elon Musk reflects on SpaceX's extremely humble beginnings, when the company operated with fewer than 10 employees and lacked even basic office furniture. This statement illustrates the minimal resour
Spent more time learning deepagents from @LangChain with primary focus on integrating MCP servers with auth that doesn't conform completely to the OAuth standard. ( In prep of a work coming my way ) B
arXiv:2602.19373v3 Announce Type: replace-cross Abstract: Deep reinforcement learning systems often suffer from unstable training dynamics due to non-stationarity, where learning objectives and data d