Extreme Self-Preference in Language Models
arXiv:2509.26464v2 Announce Type: replace Abstract: Self-preference is a fundamental feature of biological organisms. Since large language models (LLMs) lack sentience, they might be expected to avoid
Knowledge catalogue
arXiv:2509.26464v2 Announce Type: replace Abstract: Self-preference is a fundamental feature of biological organisms. Since large language models (LLMs) lack sentience, they might be expected to avoid
arXiv:2605.19859v1 Announce Type: new Abstract: Vision-language models (VLMs) have rapidly evolved into general-purpose multimodal reasoners with strong zero-shot generalization. In this context, VLMs
arXiv:2605.19291v1 Announce Type: cross Abstract: Stochastic gradient descent (SGD) is a fundamental optimization algorithm widely used in modern machine learning. In this paper, we propose Factor-Aug
arXiv:2605.19111v1 Announce Type: cross Abstract: Existing text-to-image (T2I) evaluation metrics mainly assess whether generated images align with information explicitly stated in the prompt, but oft
SpaceX filed its S-1 IPO prospectus with the SEC, offering public disclosure of its finances ahead of its investor roadshow planned for June . This filing contains the first-ever publicly disclosed in
arXiv:2605.19786v1 Announce Type: new Abstract: 4D mesh generation has recently emerged as a powerful paradigm for recovering dynamic 3D structure from videos, but existing methods remain slow, comput
arXiv:2605.19916v1 Announce Type: cross Abstract: We introduce Contrastive FUSE, a fast and unified framework for scalable node representation learning in graphs with partially available pairwise node
arXiv:2605.18908v1 Announce Type: cross Abstract: Deep neural networks (DNNs) remain critically vulnerable to backdoor attacks. Existing post-training detectors often require clean or surrogate data,
arXiv:2512.08237v3 Announce Type: replace Abstract: The advancement of vision-only Bird's-Eye-View (BEV) perception, a core paradigm for cost-effective autonomous driving, is hindered by the long-stan
arXiv:2605.19842v1 Announce Type: new Abstract: We propose a scalable tensorization framework for neural network compression based on slice-wise feature distillation. Unlike conventional tensor decomp
arXiv:2410.15362v2 Announce Type: replace-cross Abstract: Aligned Large Language Models (LLMs) have attracted significant attention for their safety, particularly in the context of jailbreak attacks t
arXiv:2509.14968v3 Announce Type: replace Abstract: The upcoming generations of wireless technologies promise an era where everything is interconnected and intelligent. As the need for intelligence gr
arXiv:2601.16200v3 Announce Type: replace-cross Abstract: Modern deep learning models exhibit strong capabilities across diverse applications, yet remain vulnerable to malicious inputs that induce err
arXiv:2605.18789v1 Announce Type: cross Abstract: Features in language models have life history: they emerge, persist, and die during training, yet the importance of that history remains largely unexp
arXiv:2601.01901v2 Announce Type: replace Abstract: Data-free knowledge distillation-based one-shot federated learning (OSFL) trains a model in a single communication round without sharing raw data, m
arXiv:2507.03122v2 Announce Type: replace-cross Abstract: This study investigates the feasibility and performance of federated learning (FL) for multi-label ICD code classification using clinical note
arXiv:2310.11203v2 Announce Type: replace Abstract: We introduce a novel strategy to train randomised predictors in federated learning, where each node of the network aims at preserving its privacy by
arXiv:2605.18936v1 Announce Type: cross Abstract: Social media text data are often used to train Machine Learning (ML) models to identify users exhibiting high-risk mental health behaviors. However, s
arXiv:2605.19949v1 Announce Type: new Abstract: Reconstructing large-scale urban scenes from sparse aerial views is a crucial yet challenging task. Due to biased top-down and shallow-oblique camera po
arXiv:2605.20016v1 Announce Type: cross Abstract: Short-form video poses new challenges to the quality assessment of user-generated content (UGC) due to its complex generation pipeline, rapid content
arXiv:2510.03589v2 Announce Type: replace Abstract: Spatio-temporal sensor data in real-world systems is often sparse, noisy, and irregular, making latent field reconstruction fundamentally underconst
Figured it out Personal update: I've joined Anthropic. I think the next few years at the frontier of LLMs will be especially formative. I am very excited to join the team here and get back to R&D. I r
arXiv:2605.20132v1 Announce Type: cross Abstract: Distributed acoustic sensing (DAS) systems generate continuous, ultra-high-channel-count data streams at rates that exceed the capabilities of convent
Financial analysts spend ~70% of their time pulling numbers out of PDFs. We built a demo agent that ingests SEC filings and answers questions with exact citations highlighted on the original PDF page.
arXiv:2602.04555v2 Announce Type: replace Abstract: Learning from a stream of tasks usually pits plasticity against stability: acquiring new knowledge often causes catastrophic forgetting of past info
arXiv:2605.18824v1 Announce Type: cross Abstract: Evaluation of foundation models often rely on aggregate scores from benchmarks that lack comprehensive coverage and metadata for a fine-grained evalua
Fine-tuned a distilbert with ml-intern today for the first time. The procedure is really straightforward, almost unexpectedly so. - Found a few datasets relevant to the task (prompt injection detectio
arXiv:2605.19224v1 Announce Type: new Abstract: Neuroscientists have recently turned to intracranial brain recording methods, like electrocorticography (ECoG), for human experiments because of the fin
arXiv:2507.10614v2 Announce Type: replace-cross Abstract: The integration of large language models (LLMs) into automated algorithm design has shown promising potential. A prevalent approach embeds LLM
arXiv:2605.20005v1 Announce Type: new Abstract: Fine-tuning large language models on new data improves task performance but degrades capabilities learned during pretraining, a phenomenon known as cata
arXiv:2605.19846v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have demonstrated remarkable capabilities in general video understanding, yet they often struggle with the fine-grained
arXiv:2509.25448v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are often modified after release through post-processing such as post-training or quantization, which makes it ch
Fireworks is coming to Tech Week First up: Boston. We're co-hosting a rooftop founder social with @fin_ai and @TrustVanta — curated crowd, real conversations, limited space. Thu May 28 · 7–10pm · #BOS
arXiv:2605.18845v1 Announce Type: cross Abstract: We give the first quantitative prediction of grokking delay under AdamW. Treating the delay as a first-passage time, we derive a closed-form law T_gro
arXiv:2605.19150v1 Announce Type: cross Abstract: State-space models (SSMs) face a fundamental trade-off between efficiency and expressivity that is mainly dictated by the structure of the model's tra
arXiv:2605.20022v1 Announce Type: new Abstract: Speculative decoding accelerates memory-bound LLM inference without quality degradation by using a fast drafter to propose multiple candidate tokens and
arXiv:2605.19739v1 Announce Type: new Abstract: Recent advances in flow matching models have significantly improved text-to-image generation quality, but also introduce growing safety risks due to the
arXiv:2605.18866v1 Announce Type: cross Abstract: Reconstructing continuous flow fields from sparse surface-mounted sensors is central to aerodynamic design, flow control, and digital-twin instrumenta
arXiv:2605.19812v1 Announce Type: cross Abstract: We introduce FLUXtrapolation, a benchmark for extrapolating ecosystem fluxes under progressively harder distribution shifts. Ecosystem fluxes are cent
arXiv:2605.19600v1 Announce Type: new Abstract: In the field of Vision-Language Navigation (VLN), aerial datasets remain limited in their ability to combine scale, diversity, and realism, often relyin
for anyone curious, this was the result of many experiments bouncing around but this version uses... initial ideation: chatgpt for ideas/@replit for quick build final repo build: opus 4.7 prompting cl
For centuries, the scientific method has been our best tool for progress. But today, there’s so much data out there that it’s impossible for any one researcher to connect all the dots. We want to fix
This entry discusses foundational concepts and context related to the Culture series, particularly the concept of the General Systems Vehicle (GSV) from Iain M. Banks' science fiction novels. The piec
arXiv:2605.19604v1 Announce Type: new Abstract: Large Language Model (LLM) agents increasingly act inside real workspaces, where tools and skills determine whether model reasoning becomes reliable act
arXiv:2605.19266v1 Announce Type: cross Abstract: Automatic speech recognition (ASR) systems are typically optimized for verbatim transcription, which preserves disfluencies, filler words, and informa
Network verification company Forward Inc. today launched Forward Predict, a new capability that lets network teams test proposed changes against a digital twin of their production network before deplo
arXiv:2605.19279v1 Announce Type: new Abstract: Visual image reconstruction from functional Magnetic Resonance Imaging (fMRI) is a fundamental task in brain decoding, providing a crucial pathway for u
France is one of the few European countries that truly understood sovereignty. It built its own nuclear deterrent, with around 290 nuclear warheads and full sovereign control over their use. It remain
arXiv:2601.18993v2 Announce Type: replace-cross Abstract: Camera redirection aims to replay a dynamic scene from a single monocular video under a user-specified camera trajectory. However, large-angle
arXiv:2605.18841v1 Announce Type: new Abstract: Safety in reinforcement learning is often specified through cumulative cost constraints, but these trajectory-level guarantees do not directly prevent u
arXiv:2605.18923v1 Announce Type: cross Abstract: Accurate selection of bovine embryos is a challenging task, as current practice relies on a single expert assessment on the seventh day after insemina
arXiv:2605.18764v1 Announce Type: cross Abstract: Artificial Intelligence (AI) pipelines have become integral to modern research, supporting fields such as Medical Sciences, Agriculture, and Social Sc
arXiv:2605.18860v1 Announce Type: cross Abstract: This paper proposes a neuron pruning framework based on neuron-level spectral structural importance evaluation. Given a trained neural network, we rec
The workforce is no longer purely human — and closing the gap between how companies manage people and how they govern AI agents has become one of the defining operational challenges of the digital wor
arXiv:2605.19824v1 Announce Type: new Abstract: Recent attempts to support high-level scene interpretation and planning in Autonomous Vehicles (AVs) using ensembles of Large Language Models (LLMs) and
arXiv:2510.13727v2 Announce Type: replace Abstract: Generative AI systems are increasingly assisting and acting on behalf of end users in practical settings, from digital shopping assistants to next-g
arXiv:2503.11615v3 Announce Type: replace Abstract: Sampling from an unknown distribution, accessible only through discrete samples, is a fundamental problem at the core of generative AI. The current
arXiv:2605.20177v1 Announce Type: new Abstract: Recent advances in vision-language models (VLMs) emphasize long chain-of-thought reasoning; yet, we find that their performance on visual tasks is prima
arXiv:2605.19781v1 Announce Type: new Abstract: Modern optimizers, like Muon, impose matrix-wise geometry constraints on their updates. These matrix-wise constraints can be unified under Linear Minimi
arXiv:2605.19263v1 Announce Type: new Abstract: Physics-informed neural networks (PINNs) offer a mesh-free framework for solving partial differential equations (PDEs), yet training often suffers from