Dr-CiK: A Testbed for Foresight-Driven Agents
arXiv:2605.27904v1 Announce Type: new Abstract: Time series forecasting in real-world settings often depends not only on historical observations, but also on external context that must be actively dis
Knowledge catalogue
arXiv:2605.27904v1 Announce Type: new Abstract: Time series forecasting in real-world settings often depends not only on historical observations, but also on external context that must be actively dis
arXiv:2605.28544v1 Announce Type: new Abstract: Pretrained foundation models have become an important basis for end-to-end autonomous driving. In contrast to vision-language models pretrained primaril
arXiv:2603.21465v2 Announce Type: replace Abstract: Developing efficient CUDA kernels is a fundamental yet challenging task in the generative AI industry. Recent research leverages Large Language Mode
arXiv:2605.27566v1 Announce Type: new Abstract: Progress in neural combinatorial optimization for Dynamic Flexible Job Shop Scheduling Problem (DFJSP) is currently hindered by a methodological tension
arXiv:2605.28510v1 Announce Type: cross Abstract: Large language models (LLMs) for code completion and generation are increasingly used in software development, yet they may reproduce training example
arXiv:2605.28573v1 Announce Type: cross Abstract: The massive scaling of Large Language Models (LLMs) has made pretraining increasingly cost-prohibitive. While low-rank representation and orthonormal
arXiv:2605.27820v1 Announce Type: new Abstract: As AI agents increasingly operate in open, real-world environments, they require a deep synergy of multimodal perception, tool invocation with multi-hop
arXiv:2605.27482v1 Announce Type: cross Abstract: While orthogonal subspace methods try to mitigate task interference in Continual Learning (CL), they often suffer from energy diffusion across the bas
arXiv:2510.27266v2 Announce Type: replace Abstract: Autonomous graphical user interface (GUI) agents rely on accurate GUI grounding, which maps language instructions to on-screen coordinates, to execu
arXiv:2605.27908v1 Announce Type: cross Abstract: Existing emotional support conversation (ESC) systems mainly rely on end-to-end response generation or coarse strategy supervision, offering limited i
arXiv:2505.17654v4 Announce Type: replace-cross Abstract: E-commerce platforms increasingly rely on Large Language Models (LLMs) and Vision Language Models (VLMs) to detect illicit or misleading produ
arXiv:2605.27618v1 Announce Type: new Abstract: Despite the wide use of explainability techniques to attempt to understand the behavior of Artificial Intelligence (AI), the generated explanations may
arXiv:2605.28312v1 Announce Type: cross Abstract: Event-based vision sensors offer asynchronous, high-temporal-resolution measurements that are attractive for low-latency robotic perception, but many
arXiv:2605.28270v1 Announce Type: new Abstract: Estimating the 9D pose of everyday objects from a single real-world image remains challenging. This is largely due to the lack of large-scale supervisio
Google created MapReduce more than 20 years ago to solve the scaling problems in data processing that the then young company was running into. The AI era that we are in now demands efficient, large-sc
arXiv:2605.27390v1 Announce Type: cross Abstract: Speculative decoding accelerates Large Language Model inference via a draft-then-verify paradigm, yet the output projection layer becomes a bottleneck
arXiv:2605.27929v1 Announce Type: cross Abstract: Active sensing links behavior and learning through an action-perception loop: actions determine the observations used to update internal predictive mo
arXiv:2604.04074v3 Announce Type: replace Abstract: LLM-based reviewing systems typically take only the manuscript as input, leaving literature and code-based claims hard to verify. We present FactRev
arXiv:2605.27486v1 Announce Type: new Abstract: Federated learning (FL) has broadened the horizon for multivariate time series anomaly detection (MTSAD). However, benchmarking such anomaly detection m
arXiv:2605.28347v1 Announce Type: new Abstract: Multi-Label Recognition (MLR) based on Vision-Language Models (VLMs) aims to leverage their pre-trained knowledge to better adapt complex recognition sc
This article likely discusses the discovery and analysis of compiler bugs or miscompilation errors in software development tools, exploring how developers identify these issues and their implications
This post discusses how compiler bugs or miscompiles can be discovered and exploited without expensive AI systems, illustrating that significant computational costs can be incurred quickly when identi
arXiv:2605.28174v1 Announce Type: cross Abstract: Foundation models offer a promising route to transferable remote sensing representations, but many current approaches depend on very large pretraining
arXiv:2605.27590v1 Announce Type: new Abstract: Remote sensing question answering (RS-QA) often requires more than direct semantic prediction, especially in large-scale forest scenes where ecological
arXiv:2605.27849v1 Announce Type: cross Abstract: Despite rapid progress in LLM-based code generation, existing models are predominantly trained on imperative languages, leaving functional programming
arXiv:2605.28188v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in high-stakes decision-making settings such as legal reasoning, where consistency under factuall
arXiv:2605.27861v1 Announce Type: cross Abstract: Predicting whether two drugs interact (binary detection) is a substantially dif- ferent task from predicting the mechanism type of that interaction (m
arXiv:2605.28303v1 Announce Type: new Abstract: While Knowledge Editing (KE) enables efficient updates, its dominant Static Fact Overwriting paradigm treats LLMs as discrete databases, forcibly inject
arXiv:2605.28359v1 Announce Type: new Abstract: Evaluating whether large language model (LLM) agents can profit in capital markets is increasingly framed as end-to-end trading: place an agent in a his
arXiv:2605.28371v1 Announce Type: new Abstract: Industrial Prognostics and Health Management (PHM) provides a representative case study for a broader challenge in applied machine learning: translating
arXiv:2605.28500v1 Announce Type: cross Abstract: Large language models have shown impressive capabilities in code generation, yet they often produce functionally incorrect code. Uncertainty quantific
arXiv:2605.28816v1 Announce Type: new Abstract: World models for interactive video generation have largely focused on single-agent settings, where future observations are generated from a single contr
glad to know Mythos' safety concerns have been addressed right as Anthropic also secured tens of billions in inference compute 👍 JUST IN: Anthropic announces it will roll out Claude Mythos “in the com
arXiv:2605.28273v1 Announce Type: new Abstract: The Policy-Space Response Oracles (PSRO) framework scales equilibrium computation to large zero-sum games by iteratively expanding a restricted strategy
arXiv:2605.27591v1 Announce Type: new Abstract: Many organizations lack computational resources to fine-tune large language models (LLMs) on private (unshareable) data for better utility, while fine-t
arXiv:2502.17055v4 Announce Type: replace-cross Abstract: Training instability in modern deep learning systems is frequently triggered by rare but extreme gradient-norm spikes, which can induce oversi
arXiv:2512.20657v2 Announce Type: replace-cross Abstract: The source detection problem arises when an epidemic process unfolds over a contact network, and the objective is to identify its point of ori
arXiv:2604.05333v3 Announce Type: replace Abstract: Modern LLM agents increasingly rely on reusable skills, and as they interact with personal applications, web browsers, and other interfaces, skill l
arXiv:2605.27595v1 Announce Type: cross Abstract: Large Language Models (LLMs) are being rapidly adopted in agricultural imaging applications, ranging from crop interpretation to synthetic field image
arXiv:2605.28315v1 Announce Type: new Abstract: General-purpose machine translation benchmarks such as FLORES-200 have reached a saturation regime on Chinese-English pairs, where modern large language
arXiv:2605.27922v1 Announce Type: new Abstract: LLM agents are increasingly deployed as executable systems that use tools, modify workspaces, and produce concrete artifacts. In such workflows, perform
arXiv:2605.28308v1 Announce Type: new Abstract: Entity Alignment (EA) is essential for knowledge graph (KG) fusion, but existing benchmarks often allow models to exploit name overlap rather than relat
This post asks Claude users for feedback about the length and depth of Claude's reasoning on tasks, soliciting examples where Claude's thinking might be excessive or insufficient. The inquiry aims to
Here Opus 4.8 built and play-tested a new RPG in Claude Code, including 3 PDF manuals and adventures, playtest notes, a website, and a playable solo adventure - then put it all on Netlify. No feedback
arXiv:2605.28554v1 Announce Type: new Abstract: Recent Tabular Foundation Models (TFMs) have demonstrated state-of-the-art predictive performance, often surpassing Gradient-Boosted Decision Trees (GBD
Holy smokes! Polymarket was not trolling. 500M accidental Claude spend in one month! Scoop from @MadisonMills22 @axios NEW: AI consultant reveals a client accidentally spent $500,000,000.00 in a singl
arXiv:2605.28302v1 Announce Type: cross Abstract: Modern large language model (LLM) inference has progressively disaggregated to keep pace with growing model sizes and tight TTFT and TPOT service-leve
'How it started, how it's going' is a popular internet meme format that compares two contrasting images or states—typically showing an initial hopeful or humble beginning alongside a current outcome t
In the high-stakes world of forensic science, time is the enemy of justice. The University of Central Oklahoma (UCO) Forensic Science Institute (FSI) was looking for an innovative AI solution that cou
https://github.com/run-llama/liteparse これか。日本語PDFでどんなか試しとこう。 We've created the world's fastest PDF parser ⚡️ And it's more accurate than any other open-source, model-free PDF parser out there (pymupdf
Mistral AI announced its participation in or perspective on the AI Now Summit 2026, likely discussing developments in AI safety, ethics, or industry trends relevant to the conference. The announcement
arXiv:2605.27724v1 Announce Type: cross Abstract: Imitation learning is a promising approach for training humanoid robots to both walk and manipulate, but it requires a large number of demonstrations,
arXiv:2605.27646v1 Announce Type: cross Abstract: We propose extbf{Hurwitz Quaternion Multiplicative Quantization (HQMQ)}, a extbf{calibration-free} method for KV cache compression of large language m
I had early access to Opus 4.8. Was impressed by it. Here is Opus 4.8's one shot of 'create a visually interesting shader that can run in twigl, make it like an infinite city of neo-gothic towers part
I had Opus 4.8 in Claude Code write a sophisticated, if minor, academic paper from a archive of hundreds of de-identified research files from years ago I had to use GPT-5.5 Pro as a reviewer, it spott
Ben Kamarov reflects on his decision to sign up for yet another SaaS product, likely discussing his evaluation criteria, the specific tool's features, or lessons learned about SaaS adoption and tool p
I think you’ll really like Opus 4.8 It’s as smart as its benchmarks show but expresses and utilizes that intelligence in a warm and collaborative way. Workflows are a great way to utilize it- I’m hook
I tried the liteparse's web browser version today to convert a couple of PDF to text and was shocked at the speed. I had to recheck twice to see whether it even did the complete processing or not 😅 ht
arXiv:2510.06928v2 Announce Type: replace Abstract: Autoregressive models have emerged as a powerful paradigm for visual content creation, but often overlook the intrinsic structural properties of vis
Connor Hart / Wall Street Journal: IBM and Red Hat commit $5B to establish a new model for open-source software, dubbed Project Lightwell, and will deploy 20,000 engineers, supported by AI — Project L