Latent Programming Horizons in Coding Agents
arXiv:2607.05188v1 Announce Type: new Abstract: A coding agent solving a software-engineering task spends dozens of steps reasoning, editing code, and running tests, yet little is known about what the
Knowledge catalogue
arXiv:2607.05188v1 Announce Type: new Abstract: A coding agent solving a software-engineering task spends dozens of steps reasoning, editing code, and running tests, yet little is known about what the
arXiv:2607.04409v1 Announce Type: new Abstract: Learning and planning in imagination using world models provides an effective paradigm for training agents for decision-making. However, existing approa
arXiv:2607.03651v1 Announce Type: new Abstract: While traditional hub capacity planning models optimize effectively for quantitative inputs, they often fail to digest qualitative business context. We
llm wikis are a glimpse of the future of what agent memory looks like this blog i wrote resonated with a lot of folks will be discussing this (as well as some updates ive made to my beliefs since this
arXiv:2607.02703v1 Announce Type: cross Abstract: In this paper, we describe LLMoxie, an institutional AI platform whose three-tiered architecture supports multi-cloud and on-premise inference, a Lite
arXiv:2607.04153v1 Announce Type: cross Abstract: Vision-based deep reinforcement learning involves dealing with high-dimensional inputs of image information. It is crucial to abstract effective state
arXiv:2607.04394v1 Announce Type: new Abstract: AI reasoning has become a central focus in contemporary artificial intelligence, largely driven by the success of large language models. However, mathem
arXiv:2607.04617v1 Announce Type: new Abstract: Long-lived AI agents require continuity across interactions, but continuity cannot be obtained by simply extending the prompt window. An agent must pres
Harrison Chase argues that evaluations (evals) are the most intellectually demanding aspect of agent engineering, distinguishing them from other engineering tasks that may be more routine or formulaic
New in Hermes Agent: pull secrets from multiple vaults at once. Run @Bitwarden and newly added vault provider @1Password side by side, or add any other secret source as a plugin with the new vault plu
arXiv:2607.05346v1 Announce Type: new Abstract: We propose OptiAgent, a multi-agent framework that, given a natural language description of an Operations Research problem, is able to output a solver-r
arXiv:2511.02200v2 Announce Type: replace Abstract: The emergence of multi-agent systems powered by large language models (LLMs) has unlocked new frontiers in complex task-solving, enabling diverse ag
arXiv:2607.03228v1 Announce Type: new Abstract: LLM-based agents offer new opportunities for automating business process execution beyond the limits of rule-based systems. However, general-purpose LLM
arXiv:2607.04576v1 Announce Type: new Abstract: LLM agents increasingly answer questions against knowledge bases they help maintain. A common intuition holds that progressive disclosure, a compact cat
arXiv:2607.02932v1 Announce Type: cross Abstract: Privacy is an important challenge when users interact with AI chatbots, since users may share sensitive information, explicitly or implicitly, and AI
Ramp 🤝 Replit Powered by Ramp, business builders on Replit can now spin up incorporation, banking, cards, and bills straight from Replit Agent. To go from idea to operating in 2026, you need agents th
arXiv:2607.04953v1 Announce Type: cross Abstract: Autonomous Driving Systems (ADS) must operate reliably under diverse conditions, yet representative data for rare or adverse scenarios is difficult to
arXiv:2607.02983v1 Announce Type: new Abstract: Recent reasoning-centric Large Language Models (LLMs) have made significant strides, yet they predominantly operate on a passive-inference pattern that
arXiv:2511.07202v3 Announce Type: replace-cross Abstract: Failures are the norm in highly complex and heterogeneous devices spanning the distributed computing continuum (DCC), from resource-constraine
arXiv:2607.03863v1 Announce Type: new Abstract: Artificial intelligence has advanced scientific discovery, but most AI4Science systems remain fragmented tools that rely on humans to coordinate problem
arXiv:2607.03616v1 Announce Type: new Abstract: As urban waste volumes escalate and labor shortages intensify, automated waste sorting systems are becoming a necessity. However, current robotic soluti
arXiv:2607.04453v1 Announce Type: cross Abstract: The assessment of planktonic standing stocks and microorganism structures is critical for understanding upper ocean biological processes. Currently, a
arXiv:2607.05382v1 Announce Type: cross Abstract: Visual generators excel at rendering, but they confidently fabricate what they do not know. User requests are unbounded, evolving, and deeply long-tai
arXiv:2607.03193v1 Announce Type: cross Abstract: Calibrating a superconducting transmon chip is a sequential decision problem under noise, drift, and a finite budget: an expert must choose experiment
arXiv:2509.10656v2 Announce Type: replace-cross Abstract: For groups of autonomous agents to achieve a particular goal, they must engage in coordination and long-horizon reasoning. Rather than relying
arXiv:2607.03726v1 Announce Type: new Abstract: While current AI agents support increasingly long context windows, tool use, and skill execution for long-horizon tasks, they still require memory syste
arXiv:2505.20203v4 Announce Type: replace Abstract: Many fear that future artificial agents will resist shutdown. I present an idea - the POST-Agents Proposal - for ensuring that doesn't happen. I pro
Since this is a slightly backwards-incompatible change here's a detailed upgrade guide - which you can read, or feed into your coding agent and have it apply the upgrades for you https://sqlite-utils.
arXiv:2607.03780v1 Announce Type: cross Abstract: SkillFab is an agent-native platform for turning missing capabilities into reviewed, reusable Agent Skills. At runtime, agents first search for reusab
arXiv:2607.02731v1 Announce Type: cross Abstract: Machine learning has demonstrated significant potential for real-time monitoring, optimization, and control of scientific facilities. However, deployi
arXiv:2607.03695v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly deployed in interacting populations, raising the question of what such populations come to believe co
arXiv:2607.05315v1 Announce Type: new Abstract: In this work, we focus on the scenario of a robot-assisted emergency evacuation. We consider two capabilities relevant to such a setting. The first is o
arXiv:2607.04732v1 Announce Type: new Abstract: Vision-based 3D semantic occupancy prediction is essential for autonomous driving, yet dense voxel representations waste computation on largely empty sp
arXiv:2607.04235v1 Announce Type: new Abstract: Large language model agents operate in partially observable, long-horizon settings where obtaining supervision remains a major bottleneck. We address th
The year 2026 could be remembered as the moment when storage technology received a massive promotion. The reason is that the current transition from simple chatbots to agentic AI systems has raised th
arXiv:2607.03216v1 Announce Type: new Abstract: Efficient flapping propulsion hinges on operating within a narrow Strouhal number window, a principle nature has converged upon for maximum thrust-to-po
arXiv:2607.04378v1 Announce Type: new Abstract: Surgical automation is being increasingly studied, yet bridging visual scene understanding with autonomous action planning remains a fundamental challen
arXiv:2607.05001v1 Announce Type: cross Abstract: Cyber Threat Intelligence (CTI) reports are predominantly unstructured, heterogeneous, and noisy, which limits their direct usability for automated an
arXiv:2503.20666v2 Announce Type: replace-cross Abstract: Thematic analysis (TA) is a widely used qualitative approach for uncovering latent meanings in unstructured text data. TA provides valuable in
arXiv:2607.04661v1 Announce Type: cross Abstract: Reconstructing 3D scene structures from sparse, low-overlap observations remains a fundamental challenge in autonomous driving. Recent state-of-the-ar
arXiv:2607.04812v1 Announce Type: new Abstract: Bird's-Eye View (BEV) end-to-end instance prediction has emerged as a robust paradigm for autonomous driving perception, effectively mitigating the erro
arXiv:2607.05168v1 Announce Type: new Abstract: Why do intelligent systems need to perform explicit symbolic reasoning? Computer science has traditionally regarded symbolic reasoning as a defining com
With the rapid progress of AI capabilities and the move to agentic systems, organizations are expanding their use cases as the technology continues to grow. That constant evolution also introduces ris
the “harness” will never go away, and it’s just as important as the model new post on harness engineering for AI self-improvement: https://lilianweng.github.io/posts/2026-07-04-harness/ It is hard to
arXiv:2607.04034v1 Announce Type: cross Abstract: The language models that underpin agents have seen a rapid rise in performance on function calling benchmarks. However, the metrics used in the traini
arXiv:2601.04387v2 Announce Type: replace Abstract: Negotiation is a core component of social intelligence, requiring agents to balance strategic reasoning, cooperation, and social norms. Recent work
This is the great thing about OpenWiki! It’ll maintain the docs for you automatically. Just setup a GitHub to run: openwiki —update and it’ll put up a PR once a day with changes / additions to your me
arXiv:2507.15903v2 Announce Type: replace-cross Abstract: Empowered by large language models (LLMs), intelligent agents have become a popular paradigm for interacting with open environments to facilit
arXiv:2607.03818v1 Announce Type: new Abstract: Indoor search-and-rescue (SAR) operations often require rapid situational awareness where GNSS signals are unavailable and human access is difficult or
Sam Schechner / Wall Street Journal: UN Secretary-General António Guterres calls for autonomous “killer robots” to be “banned by international law”, a central issue in the US DOD-Anthropic clash — Ant
arXiv:2607.05133v1 Announce Type: new Abstract: World Action Models (WAMs) have shown strong potential for improving action generalization in autonomous driving by using future video prediction as den
arXiv:2607.05277v1 Announce Type: cross Abstract: Defenses that provide security guarantees against prompt injection attacks rely on strict isolation between trusted instructions and untrusted data. I
Tim Fernholz / TechCrunch: US autonomous military vehicle startup Forterra says it deployed 100+ Lancer UGVs, based on Polaris ATVs, in Ukraine since 2025, completing 1,100+ missions — Forterra, a US
arXiv:2603.03283v2 Announce Type: replace Abstract: We dream of a future where point clouds from all domains can come together to shape a single model that benefits them all. Toward this goal, we pres
arXiv:2508.01858v3 Announce Type: replace-cross Abstract: Multimodal large-scale models have significantly advanced the development of web agents, enabling perception and interaction with digital envi
We're hosting the LA Agentic AI Meetup this Thursday, July 9, 5 to 7pm at Gulp in Playa Vista. Come talk to engineers, founders, and builders working across RAG, agentic workflows, and the broader AI
We've rolled out improvements to LlamaParse Cost Optimizer. Our intelligent tier routing now more reliably ensures you always strike the right balance between cost and accuracy when processing large d
This post likely explores the philosophy or practical approach of balancing intense work ethic with equally committed leisure and personal life, suggesting that professional ambition should be matched
yessssss Excited to announce our partnership with @tryramp to offer day-one incorporation for startups directly in Cofounder. You can now apply for incorporation, an EIN, and bank account without leav
arXiv:2607.05029v1 Announce Type: cross Abstract: Persistent memory has enabled large language model (LLM) agents to store factual knowledge, prior decisions, reasoning histories, tool usage informati