From Pixels to Prompts: Vision-Language Models
arXiv:2605.07544v1 Announce Type: new Abstract: When you read a paper about a new Vision-Language Model today, it can be easy to forget how strange this idea would have sounded not so long ago. Teachi
Knowledge catalogue
arXiv:2605.07544v1 Announce Type: new Abstract: When you read a paper about a new Vision-Language Model today, it can be easy to forget how strange this idea would have sounded not so long ago. Teachi
arXiv:2605.07816v1 Announce Type: new Abstract: This paper presents CircleID, a large-scale ICDAR 2026 competition on writer identification and pen classification from scanned hand-drawn circles. The
Wanna get a million views? Make stuff up. Take a tiny tiny bit of truth and distort it wildly. Consider the tweet below, 1.4M views. Take the chess thing. The paper that is linked doesn’t actually say
arXiv:2505.13246v2 Announce Type: replace Abstract: Purpose: This paper introduces the concept of 'Agentic Publication,' a novel LLM-driven framework designed to complement traditional scientific publ
arXiv:2605.04761v1 Announce Type: new Abstract: This paper presents the Personalized Thinking Model (PTM), a hierarchical and interpretable learner representation designed for AI supported education.
arXiv:2601.00421v2 Announce Type: replace Abstract: This paper explores how semantic-space reasoning, traditionally used in computational linguistics, can be extended to tactical decision-making in te
arXiv:2605.02010v1 Announce Type: new Abstract: This position paper argues that reliable AI requires infrastructure for human validation of implicit knowledge. AI learns from both explicit knowledge (
arXiv:2605.03443v1 Announce Type: new Abstract: This paper benchmarks classical machine learning and deep learning approaches for three-class sentiment classification of Indonesian Spotify reviews. Us
arXiv:2605.01020v1 Announce Type: new Abstract: This paper proposes and evaluates a new performance estimation method that leverages continual learning (CL) algorithms to carry out sequential simulati
arXiv:2605.01315v1 Announce Type: new Abstract: This paper investigates sentiment classification of Steam game reviews using an attention-based Bidirectional Long Short-Term Memory (BiLSTM) model. Usi
arXiv:2605.01036v1 Announce Type: new Abstract: This paper tackles the problem of physics-aware human motion synthesis in a dynamic scene. Unlike existing works which mainly tend to generate physicall
arXiv:2605.02052v1 Announce Type: new Abstract: This discussion paper reflects on how quantitative approaches to historical linguistics interact with dataset properties. Drawing on two worked examples
arXiv:2605.02212v1 Announce Type: new Abstract: This paper presents a comprehensive review of the NITRE 2026 Efficient Low Light Image Enhancement (E-LLIE) Challenge, highlighting the proposed solutio
arXiv:2603.26013v2 Announce Type: replace Abstract: Multilingual NLP is often treated as a route to global inclusion, but linguistic coverage and cultural competence frequently diverge. This paper syn
arXiv:2605.02487v1 Announce Type: new Abstract: This paper addresses the problem of mobile grasping in dynamic, unknown environments where a robot must operate under a limited field-of-view. The funda
The NeurIPS Creative AI track became part of the main conference proceedings for 2025, with papers presented as posters during the conference , marking a change from 2024 when the track was not part o
arXiv:2604.26977v1 Announce Type: cross Abstract: In response to a concern raised by Horty, this paper develops a two-tiered, preference-based semantic framework for modeling defeasible conditional ob
arXiv:2604.28158v1 Announce Type: new Abstract: Existing research infrastructure is fundamentally document-centric, providing citation links between papers but lacking explicit representations of meth
arXiv:2604.26835v1 Announce Type: cross Abstract: We introduce HalluCiteChecker, a toolkit for detecting and verifying hallucinated citations in scientific papers. While AI assistant technologies have
arXiv:2604.25267v1 Announce Type: new Abstract: This paper addresses the Dynamic UGV-UAV Cooperative Path Planning (DUCPP) problem involving one unmanned ground vehicle (UGV) assisted by one or more u
We’re excited to introduce KAME: Tandem Architecture for Enhancing Knowledge in Real-Time Speech-to-Speech Conversational AI, accepted at #ICASSP2026! 🐢 Blog https://pub.sakana.ai/kame/ Paper https://
arXiv:2604.25384v1 Announce Type: new Abstract: This paper presents a methodology for transforming raw Wikimedia dumps into quality textual corpora for seven South Slavic languages. The work is divide
arXiv:2602.17547v3 Announce Type: replace Abstract: This paper introduces KLong, an open-source LLM agent trained to solve extremely long-horizon tasks. The principle is to first cold-start the model
arXiv:2604.23733v1 Announce Type: new Abstract: Asking inquisitive questions while reading, and looking for their answers, is an important part in human discourse comprehension, curiosity, and creativ
Pay attention to this one, AI devs, especially if you're thinking about agentic commerce or any agent network where many agents share hosts. A correct route to a cold agent is still a failed request f
arXiv:2604.04353v2 Announce Type: replace-cross Abstract: Although HCI research papers offer valuable design insights, designers often struggle to apply them in design workflows due to difficulties in
arXiv:2602.21954v2 Announce Type: replace-cross Abstract: In Part I of this companion paper series, we introduced SWIFTraj, a new open-source vehicle trajectory dataset collected using a unmanned aeri
arXiv:2604.22834v1 Announce Type: new Abstract: This paper presents webmcu-vision-web, a single-file, zero-install browser application for end-to-end TinyML vision model training and deployment on the
How do AI Agents spend your money? Most teams treat agent token costs as a rounding error even though the data says they shouldn't. New paper presents the first systematic study of how agents actually
Scaling massive monolithic LLMs continues to yield incredible results. But to truly unlock their ceiling, the next frontier is test-time compute and dynamic orchestration. Nature solves complex proble
arXiv:2604.21043v1 Announce Type: cross Abstract: This paper examines the strategic use of language in contemporary artificial intelligence (AI) discourse, focusing on the widespread adoption of metap
arXiv:2604.20868v1 Announce Type: cross Abstract: In this paper, I evaluate the risks of an AI criminal mastermind, an AI agent capable of planning, coordinating, and committing a crime through the on
Tool Attention Is All You Need // Tool Attention Is All You Need // New research proposes a practical fix for the hidden 'MCP tax.' The work introduces a dynamic tool gating mechanism built on an Inte
arXiv:2604.20799v1 Announce Type: new Abstract: This paper presents a framework for mapping unknown scalar fields using a sensor-equipped autonomous robot operating in unsafe environments. The unsafe
arXiv:2604.19800v1 Announce Type: cross Abstract: This paper presents a detailed study of how graph neural networks can be used on edge intelligent meters in a microgrid to forecast photovoltaic power
arXiv:2308.00513v2 Announce Type: replace Abstract: This paper introduces UVIO, a multi-sensor framework that leverages Ultra Wide Band (UWB) technology and Visual-Inertial Odometry (VIO) to provide r
MIT researchers just replicated human muscles with AI-controlled fibers. Inside each fiber is a sealed tube of electrically charged liquid and a tiny electric pump. When the pump activates, one side c
ml-intern by @huggingface is wild 🔥 You drop a high-level prompt (“build the best scientific reasoning model” or “crush healthcare benchmarks”) and this open-source agent does the entire post-training
arXiv:2604.19645v1 Announce Type: new Abstract: An earlier paper (Hong, Potteiger, and Zapata 2026) established that an unoptimized GPT 4.1 prompt predicts fan-reported experience ratings within one p
AI Insider: 'Adding a Human Makes Your Team Worse' Emad Mostaque | @EMostaque TIMESTAMPS : 00:00 The models too dangerous to release 10:06 Why physics needs axioms — and AI doesn't 11:43 The MIND fram
arXiv:2507.12182v4 Announce Type: replace-cross Abstract: The paper is concerned with deformed Wigner random matrices. These matrices are closely related to Deep Neural Networks (DNNs): weight matrice
arXiv:2604.16865v1 Announce Type: cross Abstract: In this paper, we consider the problem of extraction of most informative features from time series that are regarded as observed values of stochastic
arXiv:2411.06812v2 Announce Type: replace-cross Abstract: This paper introduces the concept of ``generative midtended cognition'', exploring the integration of generative AI with human cognition. The
It’s time to go beyond language models. Introducing Odyssey-2 Max, our most powerful world model yet. It materially advances the SOTA in physical accuracy. This is a big step toward models that simula
arXiv:2604.16949v1 Announce Type: new Abstract: The paper considers the computation of L1 regularization paths in a state space setting, which includes L1 regularized Kalman smoothing, linear SVM, LAS
arXiv:2604.17669v1 Announce Type: new Abstract: This paper presents a comprehensive review of the NTIRE 2026 Low Light Image Enhancement Challenge, highlighting the proposed solutions and final result
arXiv:2507.16727v3 Announce Type: replace Abstract: Improving the reliability of large language models (LLMs) is critical for deploying them in real-world scenarios. In this paper, we propose extbf{De
Earlier this year Yann LeCun left Meta because Mark Zuckerberg wouldn't bet the company on JEPA. Last week his group dropped the first JEPA that actually trains end-to-end from raw pixels. 15 million
Intelligence per picojoule, with @itsclivetime and @dylan522p (0:00) Intro (1:22) What is codesign? (2:49) Codesign example: Swish vs ReLU (4:22) Are DeepSeek papers codesign? (6:45) Predicting where
arXiv:2604.14828v1 Announce Type: new Abstract: Educational assistants should spend more computation only when the task needs it. This paper rewrites our earlier draft around the system that was actua
A startup called Sabi just came out of stealth with a beanie that reads your thoughts. 70,000 to 100,000 miniature EEG sensors woven into the fabric. You put it on like a winter hat and type by imagin
arXiv:2505.19662v3 Announce Type: replace-cross Abstract: This paper introduces FieldWorkArena, a benchmark for agentic AI targeting real-world field work. With the recent increase in demand for agent
arXiv:2604.13345v1 Announce Type: new Abstract: The paper presents design and prototype implementation of an edge based object detection system within the new paradigm of AI agents orchestration. It g
This Reddit discussion thread from r/MachineLearning explores the growing challenges faced by applicants from non-elite undergraduate institutions when applying to PhD programs in machine learning and
arXiv:2604.12746v1 Announce Type: new Abstract: Stress remains a significant social problem for individuals in modern societies. This paper presents a machine learning approach for the automatic detec
arXiv:2604.09793v1 Announce Type: cross Abstract: Scientific breakthroughs often emerge from synthesizing prior ideas into novel contributions. While language models (LMs) show promise in scientific d
arXiv:2604.11261v1 Announce Type: new Abstract: This paper introduces AI as a Research Object (AI-RO), a paradigm for governing the use of generative AI in scientific research. Instead of debating whe
arXiv:2510.18976v2 Announce Type: replace Abstract: In this paper we describe Ninja Codes, neurally generated fiducial markers that can be made to naturally blend into various real-world environments.
arXiv:2604.09418v1 Announce Type: new Abstract: This paper studies Automated Instruction Revision (AIR), a rule-induction-based method for adapting large language models (LLMs) to downstream tasks usi
Introducing DDTree: accelerates speculative decoding by drafting a tree with one block diffusion pass, then verifying multiple likely continuations together. Paper: https://liranringel.github.io/ddtre