Two years ago. I stand by both predictions.
Gary Marcus reflects on predictions he made two years prior and reaffirms his confidence in their accuracy. The post suggests Marcus is reviewing his track record on forecasts, likely related to artif
Knowledge catalogue
Gary Marcus reflects on predictions he made two years prior and reaffirms his confidence in their accuracy. The post suggests Marcus is reviewing his track record on forecasts, likely related to artif
arXiv:2605.04409v1 Announce Type: new Abstract: Remote Sensing Image Change Captioning (RSICC) aims to generate spatially grounded natural language descriptions of scene evolution from bi-temporal ima
arXiv:2508.11196v2 Announce Type: replace Abstract: Recent advances in vision-language models (VLMs) have demonstrated strong generalization in natural image tasks. However, their performance often de
arXiv:2605.04941v1 Announce Type: new Abstract: This paper describes our system submitted to SemEval-2026 Task 11: Disentangling Content and Formal Reasoning in Large Language Models. We present an ef
arXiv:2511.08195v3 Announce Type: replace Abstract: UI-to-code aims to translate UI screenshots into executable front-end code. Despite progress with vision-language models (VLMs), most existing metho
arXiv:2605.04730v1 Announce Type: new Abstract: Visual localization is a core technology for augmented reality and autonomous navigation. Recent methods combine the efficient rendering of 3D Gaussian
arXiv:2605.04874v1 Announce Type: cross Abstract: Direct Preference Optimization (DPO) has proven to be an effective solution for mitigating hallucination in Multimodal Large Language Models (MLLMs) b
arXiv:2602.06869v2 Announce Type: replace Abstract: We study a persistent failure mode in multi-objective alignment for large language models (LLMs): training improves performance on only a subset of
Under-reported details of the xAI/Anthropic Colossus data center deal: Anthropic get Colossus 1 but xAI keep using the larger Colossus 2, Colossus 1 has a REALLY bad environmental record, and xAI just
arXiv:2605.05176v1 Announce Type: new Abstract: Pre-trained transformers are able to learn from examples provided as part of the prompt without any weight updates, a remarkable ability known as in-con
arXiv:2603.01097v2 Announce Type: replace Abstract: Continuous knowledge updating for pre-trained large language models (LLMs) is increasingly necessary yet remains challenging. Although inference-tim
arXiv:2508.08289v2 Announce Type: replace Abstract: Transformer architectures have revolutionized artificial intelligence (AI) through their attention mechanisms, yet the computational principles unde
arXiv:2605.04209v1 Announce Type: cross Abstract: We present Sparse Backdoor, a supply-chain attack that plants a provably undetectable backdoor in pre-trained image classifiers, including convolution
arXiv:2605.05102v1 Announce Type: new Abstract: We study the distribution of regret in stochastic multi-armed bandits and episodic reinforcement learning through a unified framework. We formalize a di
arXiv:2605.03598v2 Announce Type: cross Abstract: Understanding how biological and artificial neural networks implement computation from connectivity is a central problem in neuroscience and machine l
arXiv:2505.11815v2 Announce Type: replace Abstract: Current vision-language models have been explored for multi-modal embedding tasks like information retrieval. However, they face significant challen
arXiv:2605.04926v1 Announce Type: new Abstract: Promotional language has been increasingly used to aid the communication of innovative ideas in science. Yet, less is known about its role in the contex
arXiv:2605.04635v1 Announce Type: new Abstract: Printed Circuit Board (PCB) defect inspection faces two compounding challenges: scarce and imbalanced defect samples that limit model training, and insu
arXiv:2605.04543v1 Announce Type: new Abstract: Speculative decoding accelerates Large Language Models via draft-then-verify, where verification can be framed as an Optimal Transport (OT) problem. Exi
arXiv:2605.04819v1 Announce Type: new Abstract: Graph neural networks have been widely used in Boolean satisfiability (SAT) tasks to learn structural information from SAT formulas. The goal of these s
U.S. intelligence says Iran can outlast Trump’s Hormuz blockade for months — via @washingtonpost https://www.washingtonpost.com/national-security/2026/05/07/cia-intelligence-iran-trump-blockade-missil
arXiv:2605.04732v1 Announce Type: new Abstract: Simulation-based planning with rollouts is a widely-deployed technique for decision making in stochastic environments. The primary instrument of simulat
v0.30.0-rc3 is a release candidate version of Ollama published on May 7, 2026. As a release candidate, this version represents a pre-release milestone for the Ollama open-source project, which enables
Based on the release notes available, v0.30.0-rc4 appears to be a release candidate focused on continuous integration and platform-specific optimizations. The title suggests this build targeted Window
arXiv:2605.04078v1 Announce Type: new Abstract: Reasoning distillation aims to transfer multi-step reasoning capabilities from large language models to smaller, more efficient ones. While recent metho
arXiv:2512.05226v2 Announce Type: replace Abstract: Domain shift remains a key challenge in deploying machine learning models to the real world. Unsupervised domain adaptation (UDA) aims to address th
arXiv:2605.04750v1 Announce Type: new Abstract: Identification of less-articulated objects using single-channel images, such as thermal images, is important in many applications, such as surveillance.
arXiv:2509.14448v2 Announce Type: replace Abstract: Benchmarks such as SWE-bench and ARC-AGI demonstrate how shared datasets accelerate progress toward artificial general intelligence (AGI). We introd
arXiv:2605.04527v1 Announce Type: new Abstract: We introduce a framework for learning latent representations of 4D objects which are descriptive, faithfully capturing object geometry and appearance; c
Vercel Flags has been updated to support JSON values, expanding the types of data that can be stored and managed through the feature flags system. This enhancement allows developers to use more comple
Vibe Inc., the creator of a contextual artificial intelligence workspace platform that handles real-world meetings, introduced a wearable device today that brings an AI assistant to professionals wher
arXiv:2506.06856v3 Announce Type: replace Abstract: Visual reasoning is crucial for understanding complex multimodal data and advancing Artificial General Intelligence. Existing methods enhance the re
arXiv:2601.21851v2 Announce Type: replace Abstract: Foundation models, despite their robust zero-shot capabilities, remain vulnerable to spurious correlations and 'Clever Hans' strategies. Existing mi
arXiv:2605.04574v1 Announce Type: new Abstract: UAV-ground visual tracking (UGVT) aims to simultaneously track the same object from both the UAV and the ground view. However, existing two-stream metho
arXiv:2605.04705v1 Announce Type: cross Abstract: Today, advances in medical technology extensively utilize 3D volume data for accurate and efficient diagnostics. However, sharing these data across ne
arXiv:2605.04870v1 Announce Type: new Abstract: Video text-based visual question answering (Video TextVQA) aims to answer questions by reasoning over visual textual content appearing in videos. Despit
Vultr has launched Archival Object Storage, a new storage service offering cost-effective long-term data retention and backup solutions. This service is designed for infrequently accessed data with lo
arXiv:2605.05161v1 Announce Type: new Abstract: Zero-shot anomaly localisation via vision-language models (VLMs) offers a compelling approach for rare pathology detection, yet its performance is funda
Sam Altman argues that AI tools should be designed to augment and enhance software developers' capabilities rather than replace them, emphasizing that exceptional individual developers become extraord
We already had gemini-3.1-flash-lite-preview back on March 3rd, not clear if this new gemini-3.1-flash-lite is different other than no longer being marked as a 'preview'. Pricing appears to be the sam
We are so used to seeing chip company marketing teams exaggerate specs that it is refreshing to see them understate specs for a change. Here's one example from Cerebras's website, where they understat
We built a way to explore everything that's been entered into evidence for the Musk v. Altman trial. Greg's journal. Texts between Elon and Sam. OpenAI's LP agreement and more. All at our latest drop:
We have arrived. We are judging things based on the merit of the artist, not the technology used. You never watch a movie because how it was made or the budget it had. you watch movies because the sto
OpenAI announced on X that voice feature updates for ChatGPT are in development and coming soon, though no specific timeline was provided. The post acknowledges user demand for enhanced voice capabili
We recently found some instances of CoT grading during the training of previously deployed models after building a system that scans all OpenAI RL runs for accidental CoT grading. We did not find clea
We started running monthly research sessions at the @agi_inc office. First one was last Friday on PNM compilation. If you do on-device research and want to come present, apply here: https://form.typef
Welcome to DS4, a specialized inference engine for DeepSeek v4 Flash. https://github.com/antirez/ds4 This project would have been impossible without the existence of llama.cpp and GGML and the work of
We've teamed up with @cerebras to offer free Windsurf plans for SWE-1.6 Fast Mode at up to 1000 tok/s! Fast Mode is built on Cerebras inference, enabling superior speed for planning and development wi
Boris Cherny praised an Anthropic AI event that brought together a doctors' coding community, expressing enthusiasm about the gathering's energy and engagement. The post suggests Anthropic hosted or s
arXiv:2510.24215v4 Announce Type: replace-cross Abstract: Recovery from linear measurements under sparse adversarial corruption is typically formulated as an exact-recovery problem: one seeks structur
arXiv:2605.03354v1 Announce Type: new Abstract: Agent memory failures are silent: an LLM-based agent can produce a fluent response even when it fails to extract, retain, or retrieve the information ne
What happens when AIs become smarter than us? Why would they keep humans around if given the choice? Our new paper argues that only trying to control AIs is a limited strategy, and that a stable, mutu
arXiv:2605.05148v1 Announce Type: new Abstract: One of the major differentiators unlocked by learned codecs relative to their hard-coded traditional counterparts is their ability to be optimized direc
One of the major differentiators unlocked by learned codecs relative to their hard-coded traditional counterparts is their ability to be optimized directly to appeal to the human visual system. Despit
arXiv:2605.03782v1 Announce Type: new Abstract: To navigate partially observable visual environments, recent VLM agents increasingly internalize world modeling capabilities into their policies via exp
MIT Technology Review’s What’s Next series looks across industries, trends, and technologies to give you a first look at the future. You can read the rest of them here. Forty-eight years ago this July
arXiv:2605.03213v1 Announce Type: cross Abstract: Agentic AI systems, specifically LLM-driven agents that plan, invoke tools, maintain persistent memory, and delegate tasks to peer agents via protocol
arXiv:2605.04930v1 Announce Type: new Abstract: Despite theoretical advantages, causal methods for Gene Regulatory Network (GRN) inference from single-cell RNA-seq data consistently fail to match or o
arXiv:2507.20021v3 Announce Type: replace-cross Abstract: Recent ObjectNav systems credit large language models (LLMs) for sizable zero-shot gains, yet it remains unclear how much comes from language
arXiv:2605.05172v1 Announce Type: new Abstract: Behavior Cloning (BC) has emerged as a highly effective paradigm for robot learning. However, BC lacks a self-guided mechanism for online improvement af