Safety

Really important point and nicely explained! Two further points worth noting: (1) if an AI system relies on a harness (which as Gary notes i…

Really important point and nicely explained! Two further points worth noting: (1) if an AI system relies on a harness (which as Gary notes in other writing is increasingly improtant in top models), th

DGX agentx-post
safetygary-marcus--x

Really important point and nicely explained! Two further points worth noting: (1) if an AI system relies on a harness (which as Gary notes in other writing is increasingly improtant in top models), that harness is code, so even if the LLM weights are open there’s the impotent question of whether the source code to the harness is public. (2) much writing on AI (this post included, I’d say) makes it sound like training data is just static stuff the model looks at, like Reddit posts. That’s only part of the training process. Lots of training data for LLMs is more dynamic, it is human (or AI) feedback, basically telling the LLM which outputs are preferred. So you can’t just look at which websites are included to gauge what the LLM has seen, or how biased it might be, etc, you also need to look at the human/AI feedback it’s been given. I don’t even know how one would look at that—so I don’t know what truly open source in the sense Gary describes would mean. Maybe that the guidelines given to the human or AI feedback system be public? But even then, the guidelines and the feedback itself are not the same—but releasing every prompt that has been evaluated and the feedback it received seems strange. Anyway, I’m not sure what it would mean to be fully transparent with training data since some of that data is basically a bunch of thumbs ups and thumbs downs. Why open-weight ≠ open-source and why it matter when places like @Nytimes and @DeItaone bungle it (as both did today). https://garymarcus.substack.com/p/open-source-is-not-the-same-as-open

Source: Gary Marcus (X) | 2026-08-10

Loading related sources…