self-modification
Self-modification in AI systems refers to the capability of an artificial intelligence to alter its own code, parameters, or behavior patterns without external intervention. This concept, discussed by
Knowledge catalogue
Self-modification in AI systems refers to the capability of an artificial intelligence to alter its own code, parameters, or behavior patterns without external intervention. This concept, discussed by
arXiv:2605.29558v1 Announce Type: new Abstract: Severe image degradation under low-light nighttime conditions constitutes a core bottleneck preventing all-day applications for UAV-based single object
arXiv:2605.30151v1 Announce Type: new Abstract: As AI tools become increasingly integrated into educational contexts, questions arise about both their stability over time and their responsiveness to p
Tip on Grok Build + Grok 4.3 VLM One of my critical tasks is to keep Grok VLM in the loop. Throwing a default system prompt usually yields poor results due to lack of context. Here is how to scale: -
65B private round More than double the size of the largest IPO ever We've raised 65 billion in Series H funding at a $965 billion post-money valuation, led by @AltimeterCap, Dragoneer, @Greenoaks, and
arXiv:2605.27873v1 Announce Type: new Abstract: AI models underpin data-centric applications from image and text processing to scientific discovery in biology, physics, and chemistry. Yet developing t
Cognition, an AI company, raised 1 billion in Series D funding at a 26 billion valuation, indicating significant investor confidence in its technology and business model. The funding round reflects th
arXiv:2605.28642v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have demonstrated significant potential for speech-to-text translation (S2TT). However, existing deployment par
arXiv:2605.27407v1 Announce Type: cross Abstract: Evaluating fairness in Spiking Neural Networks (SNNs) demands rigorous benchmarks that reflect real-world complexities, yet existing assessments remai
BREAKING: Anthropic just dropped Opus 4.8—and it is a MONSTER We've been testing for about a week @every and our verdict is they could've just called it Opus 5, it's that good. Here's our vibe check:
arXiv:2605.28257v1 Announce Type: new Abstract: Understanding 3D objects from images is fundamental to robotics and AR/VR applications. While recent work has made progress in category-level pose estim
arXiv:2603.22335v2 Announce Type: replace-cross Abstract: Direct Preference Optimization (DPO) guides large language models (LLMs) to generate recommendations aligned with user historical behavior dis
arXiv:2605.27449v1 Announce Type: cross Abstract: In the field of multimodal fact checking, the accuracy of retrieving evidence from different modalities has a significant impact on the downstream cla
Anthropic shipped Claude Opus 4.8 today. My favourite thing about it is this note in the release announcement: Users will find Opus 4.8 to be a modest but tangible improvement on its predecessor. Ther
arXiv:2512.00349v2 Announce Type: replace Abstract: Are frontier AI systems becoming more capable? Certainly. Yet such progress is not an unalloyed blessing but rather a Trojan horse: behind their per
arXiv:2605.28104v1 Announce Type: new Abstract: Recent years have witnessed the rapid development of Large Language Model-based Multi-Agent Systems (MAS), which excel at collaborative decision-making
arXiv:2605.28148v1 Announce Type: cross Abstract: The rapid development of LLMs coupled with the introduction of Model Context Protocol (MCP) has revolutionized how intelligent agents interact with AP
Cal Newport discusses whether recent AI advances represent a breakthrough in mathematical problem-solving capabilities, examining the implications of AI systems' improved performance on complex mathem
arXiv:2605.27935v1 Announce Type: new Abstract: Recent mechanistic studies suggest that large language models (LLMs) may utilize their depth inefficiently in standard single-turn tasks. Whether this s
arXiv:2605.27486v1 Announce Type: new Abstract: Federated learning (FL) has broadened the horizon for multivariate time series anomaly detection (MTSAD). However, benchmarking such anomaly detection m
This post discusses how compiler bugs or miscompiles can be discovered and exploited without expensive AI systems, illustrating that significant computational costs can be incurred quickly when identi
Gary Marcus shares a discussion thread about consciousness with user @Grimezsz on X (formerly Twitter), likely exploring philosophical, scientific, or technical perspectives on consciousness and relat
arXiv:2512.20657v2 Announce Type: replace-cross Abstract: The source detection problem arises when an epidemic process unfolds over a contact network, and the objective is to identify its point of ori
arXiv:2605.27433v1 Announce Type: cross Abstract: With the increasing complexity of collaboration among various social entities and user demands, the factors affecting the stable development of the da
arXiv:2605.27373v1 Announce Type: new Abstract: As intelligent systems become more autonomous, the scientific community focuses on creating decision-making mechanisms that include ethical and moral co
if you replace billions with millions, this sounds like any other high-growth startup fundraise announcement 😉 We've raised 65 billion in Series H funding at a 965 billion post-money valuation, led by
arXiv:2605.27989v1 Announce Type: new Abstract: The guidance of scaling laws has increased the resource demands of modern large language models (LLMs), yet it remains questionable whether these models
arXiv:2605.27391v1 Announce Type: cross Abstract: This paper examines whether students are entering the generative AI era with sufficiently strong educational foundations, focusing on the relationship
arXiv:2605.28120v1 Announce Type: cross Abstract: Graph-based Retrieval-Augmented Generation (GraphRAG) advances flat document retrieval by structuring knowledge as relational graphs, enabling more co
arXiv:2510.01724v2 Announce Type: replace Abstract: Mass spectrometry-based metabolomics generates complex, high-dimensional data that holds vast potential for biological discovery but remains difficu
arXiv:2502.17832v4 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) has become a common practice in multimodal large language models (MLLM) to enhance factual grounding and
arXiv:2605.28137v1 Announce Type: new Abstract: Text-to-image models trained on large-scale data often inevitably ingest unsafe content. While some people observe input-output amplifications, it remai
arXiv:2605.28168v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated promising capability in generating reward functions for deep reinforcement learning (DRL)-based building
arXiv:2601.15015v2 Announce Type: replace Abstract: Reinforcement learning (RL) has shown promising results in active flow control (AFC), yet progress in the field remains difficult to assess as exist
arXiv:2605.28241v1 Announce Type: new Abstract: Point cloud quality plays a critical role in 3D acquisition, reconstruction, rendering, and perception, yet existing point cloud quality assessment (PCQ
arXiv:2605.28375v1 Announce Type: new Abstract: Prion diseases are rare, rapidly progressive, and fatal neurodegenerative disorders that remain difficult to diagnose, particularly in their early stage
📢Qwen3.7-Max just hit #3 on ITbench-AA — a fresh benchmark testing how well models handle real-world enterprise IT tasks, agentic-style. 🔧Agentic era, go with Qwen.🏃🏃 Artificial Analysis and IBM Resea
arXiv:2502.08938v4 Announce Type: replace Abstract: In the past decade, motivated by the putative failure of naive self-play deep reinforcement learning (DRL) in adversarial imperfect-information game
arXiv:2512.12887v3 Announce Type: replace Abstract: 3D medical image classification is essential for modern clinical workflows. Medical foundation models (FMs) have emerged as a promising approach for
arXiv:2605.27911v1 Announce Type: new Abstract: Suicide is a critical global public health challenge, causing approximately 720,000 deaths each year and calling for timely, effective prevention strate
Super excited to finally share Dynamic Workflows in Claude Code!! We built this a couple months ago, and it has slowly become a daily driver for a bunch of people at Anthropic. A few tips for getting
arXiv:2407.12014v1 Announce Type: cross Abstract: Autism is a developmental disorder that manifests in early childhood and persists throughout life, profoundly affecting social behavior and hindering
arXiv:2605.27901v1 Announce Type: cross Abstract: Chain-of-thought (CoT) monitoring has been proposed as a promising safety mechanism for detecting misaligned behavior in large language models. Howeve
The HF science team just made async RL weight sync ~100x cheaper on bandwidth, and you don't need a shared cluster anymore. The problem: every RL step, the trainer typically has to sync fresh weights
arXiv:2605.28102v1 Announce Type: new Abstract: Large language models trained with Reinforcement Learning from Human Feedback (RLHF) and Constitutional AI exhibit persistent behavioral patterns that s
The importance of Snowflake Inc. in the enterprise AI ecosystem is not necessarily its contribution to the analytics warehouse. What will be more significant going forward is Snowflake’s role as an en
arXiv:2605.27580v1 Announce Type: new Abstract: A central puzzle for the behavioural sciences and for human-facing artificial intelligence is the persistence of within-person variability. The same ind
arXiv:2605.26747v1 Announce Type: new Abstract: Large Language Models (LLMs) have brought huge improvements to Artificial Intelligence (AI), which can be applied to general-purpose tasks. However, the
I saw a developer asking on Reddit if there was any “sane way” to manage Cloud Run cold starts for AI across multiple regions. They were experiencing startup latencies of up to 20 seconds, a frustrati
DeepMind's AlphaFold3 and related tools have generated a comprehensive atlas containing over one billion predicted protein structures and additional billions of protein sequences, representing a major
arXiv:2511.02525v2 Announce Type: replace-cross Abstract: The capacitated location-routing problems (CLRPs) are classical problems in combinatorial optimization, which require simultaneously making lo
arXiv:2503.21510v3 Announce Type: replace-cross Abstract: Ensuring that predictions of machine learning (ML) classification models are accompanied by uncertainty estimates is one of the main pillars o
arXiv:2605.26397v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used in decision-making tasks where they can amplify or suppress perspectives, raising concerns in high-
ANNOUNCEMENT: WE’RE SAVING SCIENCE! We’re often told that science is “self-correcting.” But that’s not really true. Science doesn’t correct itself like a thermostat adjusting the temperature in your h
arXiv:2605.27050v1 Announce Type: new Abstract: We present BhashaSetu, a linguistically enriched English--Marathi parallel dataset addressing persistent data limitations in low-resource neural machine
arXiv:2605.26820v1 Announce Type: new Abstract: Vision-language-action (VLA) models provide a promising foundation for general-purpose robotics. However, their successful deployment in real-world scen
arXiv:2605.26675v1 Announce Type: cross Abstract: CART random forests are among the most widely used modern predictive methods, with well-documented empirical success. Yet, at the mechanistic level, t
arXiv:2605.26774v1 Announce Type: new Abstract: Cesarean Scar Defect (CSD) is one of the most prevalent complications following cesarean delivery. Transvaginal ultrasonography is widely used for prima
arXiv:2605.26734v1 Announce Type: new Abstract: Existing Multi-Turn Composed Image Retrieval (MTCIR) datasets lack dialogue-history consistency and are restricted to the fashion domain. To address the
A new report out today from Cisco Systems Inc. argues that none of the closed flagship large language models it tested can be considered safe once an attacker is allowed to push past a single prompt,