Model Releases

ran my first benchmark this weekend (longmemeval) mostly to test activegraph, learned a lot! - this is a stepping stone to show the event ba…

ran my first benchmark this weekend (longmemeval) mostly to test activegraph, learned a lot! - this is a stepping stone to show the event based agent system works. the AI convinced me not to start wit

DGX agentx-post
model-releasesyohei-nakajima--x

ran my first benchmark this weekend (longmemeval) mostly to test activegraph, learned a lot! - this is a stepping stone to show the event based agent system works. the AI convinced me not to start with graph extraction of facts/entities - learned running benchmarks takes a long time - understand more what a good benchmark vs bad benchmark is (I think this is pretty thorough) - fully reproducible open source repo and tests - seems like ActiveGraph is well suited for this, performed solid, which is a good start [technical blog post] “Evidence Compilation Before Semantic Memory: ActiveGraph on LongMemEval-S” —— 🔍85.6% QA accuracy and 86.2% turn answer-in-context at 2,462 mean context tokens, with deterministic non-generative ingestion https://activegraph.ai/blog/evidence-compilation-bef…

Source: Yohei Nakajima (X) | 2026-05-26

Loading related sources…