Automated conjecturing with TxGraffiti
arXiv:2409.19379v2 Announce Type: replace-cross Abstract: TxGraffiti is a data-driven, heuristic-based computer program developed to automate the process of generating conjectures across various mathe
Knowledge catalogue
arXiv:2409.19379v2 Announce Type: replace-cross Abstract: TxGraffiti is a data-driven, heuristic-based computer program developed to automate the process of generating conjectures across various mathe
arXiv:2605.10370v1 Announce Type: new Abstract: Scientific knowledge on the Web is published as passive assertions and cannot decide when to validate evidence, reconcile contradictions, or update conf
arXiv:2605.08110v1 Announce Type: cross Abstract: Low-Rank Adaptation (LoRA) has become the standard for fine-tuning large pre-trained models at reduced computational cost. However, its low-rank point
arXiv:2605.08214v1 Announce Type: cross Abstract: Automatic Speech Recognition (ASR) and speaker diarization in Bangla remain challenging due to long form recordings, diverse acoustic conditions, and
arXiv:2510.09877v3 Announce Type: replace-cross Abstract: Over the past couple of decades, many active learning acquisition functions have been proposed, leaving practitioners with an unclear choice o
arXiv:2601.02950v3 Announce Type: replace Abstract: Current Large Language Model reasoning systems process queries independently, discarding valuable cross-instance signals such as shared reasoning pa
arXiv:2605.10867v1 Announce Type: cross Abstract: Continuous authentication in high-stakes digital environments requires datasets with fine-grained behavioral signals under realistic cognitive and mot
arXiv:2605.08463v1 Announce Type: new Abstract: Autonomous AI agents are increasingly deployed in open social environments, yet the relationship between their configuration specifications and their em
arXiv:2605.08405v1 Announce Type: new Abstract: How do LLMs learn in-context? Is it by pattern-matching recent tokens, or by inferring latent structure? We probe this question using a toy graph random
arXiv:2605.10865v1 Announce Type: new Abstract: Industrial Computer-Aided Design (CAD) code generation requires models to produce executable parametric programs from visual or textual inputs. Beyond r
arXiv:2605.08988v1 Announce Type: cross Abstract: Machine Learning Interatomic Potentials play a fundamental role in computational chemistry and materials science, enabling applications from molecular
arXiv:2605.08136v1 Announce Type: cross Abstract: Visual perception plays a central role in competitive robotics, where environmental variations can directly affect real-time detection performance. Th
arXiv:2605.10146v1 Announce Type: new Abstract: Large language models (LLMs) increasingly rely on knowledge editing to support knowledge-intensive reasoning, but this flexibility also introduces criti
arXiv:2407.12173v2 Announce Type: replace-cross Abstract: Generative diffusion models have emerged as a powerful tool for high-quality image synthesis, yet their iterative nature demands significant c
arXiv:2605.09292v1 Announce Type: new Abstract: Large language models now achieve high final-answer accuracy on mathematical reasoning benchmarks, but accuracy alone does not capture reasoning flexibi
arXiv:2605.10223v1 Announce Type: new Abstract: Current large language model agent frameworks prioritize autonomy but lack the governability mechanisms required for enterprise deployment. High-risk wr
arXiv:2605.09268v1 Announce Type: cross Abstract: Users interacting with Large Language Models (LLMs) in a multi-turn conversation routinely refine their requests or pivot to new topics. LLMs, however
arXiv:2605.09310v1 Announce Type: new Abstract: ESG-aware portfolio optimization is increasingly important for sustainable capital allocation, yet most learning-based methods still operationalize ESG
arXiv:2605.08202v1 Announce Type: cross Abstract: Offline reinforcement learning (RL) faces a critical challenge of overestimating the value of out-of-distribution (OOD) actions. Existing methods miti
arXiv:2605.09153v1 Announce Type: cross Abstract: Closed-loop traffic simulation requires agents that are both scalable and behaviorally realistic. Recent self-play reinforcement learning approaches d
arXiv:2605.08280v1 Announce Type: cross Abstract: Preserving model fidelity is essential for stealthy text-to-image (T2I) backdoor attacks. Existing methods such as Learning without Forgetting (LwF) r
arXiv:2605.10780v1 Announce Type: cross Abstract: Representation autoencoders that reuse frozen pretrained vision encoders as visual tokenizers have achieved strong reconstruction and generation quali
arXiv:2502.08943v4 Announce Type: replace-cross Abstract: Large language models (LLMs) have demonstrated significant utility in real-world applications, exhibiting impressive capabilities in natural l
arXiv:2605.08716v1 Announce Type: new Abstract: Are certain cognitive biases mathematically inevitable consequences of sequential information processing? We prove that primacy effects, anchoring, and
arXiv:2512.24863v2 Announce Type: replace-cross Abstract: The world is in the grip of ecological, meaning, and language crises that are converging into a metacrisis. Big AI is accelerating them all. L
arXiv:2605.08564v1 Announce Type: new Abstract: The feedback alignment (FA) algorithm offers a biologically plausible alternative to backpropagation (BP) for training neural networks yet notably fails
arXiv:2605.09579v1 Announce Type: cross Abstract: Cardiovascular disease remains the leading cause of global mortality, yet scalable cardiac monitoring is hindered by the gap between diagnostic-rich E
arXiv:2602.07144v2 Announce Type: replace-cross Abstract: Bayesian optimization (BO) is a popular technique for sample-efficient optimization of black-box functions. In many applications, the paramete
arXiv:2605.09134v1 Announce Type: new Abstract: Reinforcement learning for program repair is hindered by sparse execution feedback and coarse sequence-level rewards that obscure which edits actually f
arXiv:2605.10764v1 Announce Type: cross Abstract: Recent studies show that gradient-based universal image jailbreaks on vision-language models (VLMs) exhibit little or no cross-model transferability,
arXiv:2603.04783v2 Announce Type: replace Abstract: While LLMs demonstrate strong reasoning capabilities when provided with full information in a single turn, they exhibit substantial vulnerability in
arXiv:2602.08616v2 Announce Type: replace-cross Abstract: Reinforcement Learning (RL) is increasingly applied to large-scale decision-making problems like logistics, scheduling, and recommender system
arXiv:2605.08271v1 Announce Type: cross Abstract: Understanding ultra-long videos such as egocentric recordings, live streams, or surveillance footage spanning days to weeks, remains a challenge. For
arXiv:2605.10541v1 Announce Type: new Abstract: Epigenetic clocks based on DNA methylation have emerged as powerful tools for estimating biological age, with broad applications in aging research, age-
arXiv:2605.10036v1 Announce Type: cross Abstract: As 6G evolves, the radio access network must transcend traditional automation to embrace agentic AI capable of perception, reasoning, and evolution. A
arXiv:2605.08862v1 Announce Type: cross Abstract: Reinforcement Learning (RL) has become a cornerstone for improving the performance of Large Language Models (LLMs). However, its rollout phase constit
arXiv:2605.10598v1 Announce Type: new Abstract: Large language models (LLMs) have emerged as powerful tools for automatic algorithm design (AAD). However, existing pipelines remain inefficient. They o
arXiv:2605.08404v1 Announce Type: cross Abstract: This work investigates the use of large language models (LLMs) for tasks in smart cities. The core idea is to leverage remote sensing imagery to chara
arXiv:2605.10661v1 Announce Type: cross Abstract: Vision Transformers (ViTs) are built by stacking independently parameterized blocks, but it remains unclear how much of this depth requires layer spec
arXiv:2605.08653v1 Announce Type: new Abstract: Accurate state-of-charge (SOC) estimation is critical for the safe and efficient operation of lithium-ion batteries in battery management systems (BMS).
arXiv:2504.21228v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are susceptible to indirect prompt injection attacks, where the model inadvertently responds to instructions inje
arXiv:2605.10873v1 Announce Type: cross Abstract: Recovering editable CAD programs from images or 3D observations is central to AI-assisted design, but progress is difficult to measure because existin
arXiv:2605.09823v1 Announce Type: cross Abstract: We introduce CalBench, a controlled evaluation environment for studying multi-agent coordination through calendar scheduling. In CalBench, N agents ea
arXiv:2605.08325v1 Announce Type: cross Abstract: Many vision datasets now provide segmentation masks in addition to annotated images to support a wide range of tasks. In this work, we propose Class A
arXiv:2605.10448v1 Announce Type: new Abstract: Interactive agent benchmarks map an agent run to a binary outcome through outcome checks. When these checks rely on surface level signals or fail to cap
arXiv:2605.10419v1 Announce Type: cross Abstract: This paper investigates the effectiveness of large language models (LLMs) in answering questions over datasets. We examine their performance in two sc
arXiv:2512.18880v2 Announce Type: replace-cross Abstract: Accurate estimation of item (question or task) difficulty is critical for educational assessment but suffers from the cold start problem. Whil
arXiv:2605.08255v1 Announce Type: cross Abstract: Can large language models predict physical and mechanical polymer properties simply by reading unstructured scientific prose? Polymer performance is r
arXiv:2605.06638v2 Announce Type: replace Abstract: Reinforcement learning (RL) has been applied to improve large language model (LLM) reasoning, yet the systematic study of how training scales with t
arXiv:2605.08938v1 Announce Type: new Abstract: Fourier Neural Operators (FNOs) can greatly accelerate PDE simulation, but they are often used without formal guarantees that they preserve basic physic
arXiv:2605.10794v1 Announce Type: cross Abstract: Language models are deployed in settings that require compartmentalization: system prompts should not be disclosed, chain-of-thought reasoning is hidd
arXiv:2503.05066v5 Announce Type: replace-cross Abstract: The Mixture of Experts (MoE) is an effective architecture for scaling large language models by leveraging sparse expert activation to balance
arXiv:2512.04949v3 Announce Type: replace-cross Abstract: Agents capable of accomplishing complex tasks through multiple interactions with the environment have emerged as a popular research direction.
arXiv:2605.09016v1 Announce Type: new Abstract: Neural operators have emerged as powerful data-driven solvers for PDEs, offering substantial acceleration over classical numerical methods. However, exi
arXiv:2605.08740v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) decompose transformer residual streams into interpretable feature dictionaries, yet the relationship between SAE width and
arXiv:2507.03310v2 Announce Type: replace-cross Abstract: This paper studies causal discovery in irregularly sampled time series-a key challenge in risk-sensitive domains like finance, healthcare, and
arXiv:2605.09663v1 Announce Type: cross Abstract: Machine learning classifiers in dynamic environments face concept drift -- changes in the data-generating process that degrade performance. Convention
arXiv:2605.08590v1 Announce Type: cross Abstract: LLMs are increasingly used to explain personal sensing data, translating traces of activity and mood into natural-language accounts of why an anomalou
arXiv:2605.09079v1 Announce Type: new Abstract: Despite surpassing human performance across mathematics, coding, and other knowledge-intensive tasks, large language models (LLMs) continue to struggle
arXiv:2605.10473v1 Announce Type: cross Abstract: We introduce a cavity-enhanced optical architecture for collective quantum processing in which logical qubits are encoded in the polarization subspace