Model Releases
Two major shifts will be seen in Agentic AI after Harness and YOU MUST KNOW. 1. Workflow design of your agents matters a lot more than any f…
Two major shifts will be seen in Agentic AI after Harness and YOU MUST KNOW. 1. Workflow design of your agents matters a lot more than any frontier model selection. Till now we have mostly focused on
Two major shifts will be seen in Agentic AI after Harness and YOU MUST KNOW. 1. Workflow design of your agents matters a lot more than any frontier model selection. Till now we have mostly focused on chasing leaderboard models and burning money on frontier models. LLMs have a decoding architecture—they just predict the next words and give answers based on what we fed them while pre-training. That's why we started implementing RAG and multi-agent tool calling. This improves output but we lose control of quality and precise output. Feeding knowledge leads to context problems. That's where context engineering comes into the picture. Make your prompt/input much cleaner and keep only important information to get quality output. Remember, Garbage in → Garbage Out Moving towards Harness, it has a cleaner design and goes beyond context engineering. It focuses more on tool outputs, memory access and control, additionally LLM as a Judge—which decides what to pick, what to send as input, and what to keep for output. Just like machine learning, where your system matters a lot more than just accuracy. Exactly, we need to focus more on orchestration workflow of agents than just the model. --------------------- 2. No place for hallucination as memory kicks in: In 2026, you will experience hallucination in chain of conversation, not at the very first question. LLMs are getting smarter and have external tool access, especially web search. But as conversation grows, you stop sending detailed prompts and LLMs cannot process them. Example: There are two Ronaldos, CR7 and R9. If you are talking about CR7 for a long time and suddenly ask, "how was his performance in 2002 WC?" the LLM gets confused as CR7 never played 2002 WC. Now you are indirectly asking about R9. So you need memory switching in the conversation. Short-term and long-term memories are really game-changing. You can easily get better output on confusing questions if you switch between memories. Most LLM providers lock your memory and it's not worth staying relevant in longer, confusing conversations. Short memory is more like a cache that picks your latest conversations, and long-term acts like a database but in summarized format. Keep only what matters for longer memory. Time-aware conversation is also very important. Ask Claude about what you asked on 5-11 Dec 2025, it wasn't able to tell you because there's no time-aware conversation. I've seen most people are working on that. If that comes into the picture, you will see real intelligence. It's like someone remembering everything with date-time. @hwchase17 explained more about Harness and memory importance in this article. He also talked about Deep Agents which might be a gamechanger again. Lastly, How you design your system architecture from a user perspective matters more than Opus and GPT.
Source: Harrison Chase (X) | 2026-04-14