Agents

longmemeval experiment arch: 1) deterministic ingestion/extraction (85.6% accuracy, 86.2% retrieval) 2) semantic ingestion/extraction (84.8%…

longmemeval experiment arch: 1) deterministic ingestion/extraction (85.6% accuracy, 86.2% retrieval) 2) semantic ingestion/extraction (84.8% accuracy, 94.9% retention) 3) semantic ingestion/determinis

DGX agentx-post
agentsyohei-nakajima--x

longmemeval experiment arch: 1) deterministic ingestion/extraction (85.6% accuracy, 86.2% retrieval) 2) semantic ingestion/extraction (84.8% accuracy, 94.9% retention) 3) semantic ingestion/deterministic extraction (87.6% accuracy, 90.0% retrieval) this was staggered, not linear, so i did not build each experiment on top of the other so they are not apple to apple comparisons but the real purpose here was to showcase the benefit of experimenting auditable agents Our third LongMemEval experiments by adding semantic ingestion to deterministic retrieval. The hybrid approach improved end-to-end QA accuracy to 87.6% and evidence retrieval to 90.0%, but the QA gain did not reach statistical significance in this run. https://activegraph.ai/blog…

Source: Yohei Nakajima (X) | 2026-06-01

Loading related sources…