I spent the summer of 2017 hunched over solidity code in a Seattle basement, checking for reentrancy bugs in ICO smart contracts. Three projects had vulnerabilities that could have drained user funds. The founders meant well—they were just copying incentive structures without understanding the game theory behind them. This week, as I read the parsed analysis of Moonshot AI founder Yang Zhilin's interview on using Reinforcement Learning (RL) and Supervised Fine-Tuning (SFT) as management analogies, I felt the same tension. A compelling narrative, but one that unravels when you map it to the messy reality of human coordination—something I’ve been mapping in liquidity flows for years.
Listening to the silence between market cycles, I notice how often new management philosophies borrow from technical paradigms without addressing the alignment failures that made those paradigms necessary in the first place. Yang’s analogy is elegant: RL as a management style that lets employees explore freely with rewards; SFT as direct instruction via labeled data. It fits the current AI company culture playbook—autonomy, innovation, high talent density. But as someone who has watched DeFi summer yield farms collapse when their reward functions were gamed, I see the unspoken risks.
The context here matters beyond corporate culture. Moonshot AI (the company behind Kimi) operates in a space where the boundaries between AI agents and crypto agents are blurring. In 2026, I published a study on AI-crypto symbiosis, analyzing 50,000 automated transactions. The convergence demands governance models that are transparent, auditable, and resistant to reward hacking. Yang’s RL analogy, while thought-provoking, simplifies the core challenge: how do you design a reward function in human organizations that aligns with long-term value, not short-term metrics? In crypto, we see this daily—liquidity mining programs that pump TVL but attract mercenary capital. The RL management approach risks the same: employees optimizing for the visible reward (code commits, project milestones) while neglecting the invisible (code maintainability, team cohesion).
Based on my DeFi summer liquidity mapping experience, I traced how Federal Reserve liquidity injections flowed into Uniswap pools, creating temporary arbitrage opportunities. The behavior was rational per the reward structure but collapsed months later. Yang’s analogy misses the temporal alignment problem: RL in AI works because the environment is controlled; in human organizations, the reward function must evolve with context. Without explicit guardrails—what I call a ‘management constitution’ similar to constitutional AI—the RL approach can breed a toxic competition where employees learn to game the system.
Here is where the contrarian angle emerges. The article analysis correctly notes that missing technical depth—sparse rewards, credit assignment, multi-agent coordination. But the deeper insight is that pure RL management is actually a step backward from the decentralized coordination possibilities that crypto offers. Smart contracts provide a transparent, immutable reward function. They remove the human bias and opacity that make RL management risky. Yang’s model is still top-down: the manager defines the reward proxy. In contrast, on-chain governance, like that in DAOs, allows reward functions to be voted on and adjusted programmatically. The future of AI company governance might not be RL or SFT but hybrid systems where base rules are encoded in smart contracts (SFT-like) and exploration is incentivized via tokens (RL-like). I saw this potential during the 2022 bear market when we hosted webinars on trust and verification—people craved transparent mechanisms, not charismatic founder analogies.
Listening to the silence between market cycles, I hear the noise of yet another management fad borrowing from AI. But the real innovation is not in replicating RL for humans; it is in using crypto’s toolset to design governance that is self-correcting. Yang’s interview is a signal that AI leaders are thinking about coordination, but they are overlooking the most radical solution: decentralized, algorithmic trust. My 2024 ETF study showed that institutional capital demands transparency; the same will apply to AI company governance as they scale. The question is not whether RL management works for a 100-person team, but whether it can scale without the accountability that on-chain systems provide.
As I finish this article, I recall the 2017 ICO audits. The founders who failed were not malicious; they just lacked the frameworks to foresee unintended consequences. Yang’s analogy is thoughtful, but it is still a central planning approach dressed in technical metaphors. The field needs to look beyond management philosophy to the infrastructure of trust itself. What if Moonshot AI published its reward function as a smart contract, letting the community audit it? That would be a true decentralization play.
Listening to the silence between market cycles, I wait for the moment when the industry stops analogizing and starts building the verifiable coordination layers that will define the next decade. The RL management mirage is a distraction—the real frontier is programmable governance.


