The RL vs SFT Battle for Crypto Teams: Why Moonshot's Management Metaphor Reveals a Deeper Truth About Token Incentives

SamWolf Price Analysis
Hook Yang Zhilin, founder of Moonshot AI, recently dropped a management metaphor that rippled through the crypto-brain trust: treat your team like a reinforcement learning (RL) agent, not a supervised fine-tuning (SFT) one. Let them explore. Set rewards, not instructions. The interview was framed as a leadership insight, but for anyone who has watched a DeFi protocol bleed LPs after a bad tokenomic tweak, it hits closer to home. RL and SFT aren't just training paradigms—they are the two poles of every crypto project's incentive design. And most teams are building on the wrong one. Context The analogy is elegant on the surface. In AI, SFT means feeding a model labeled examples: “do this, not that.” RL means dropping it into an environment with a reward function and letting it discover optimal behavior through trial and error. Yang argued that modern teams need RL-dominant cultures—autonomy, experimentation, outcome-driven rather than rule-driven. The crypto space has lived both extremes. The ICO era was pure RL: founders threw tokens into the wild and let the community game the system. The 2021 NFT bull run was SFT on steroids: floor price targets, curated whitelists, step-by-step roadmaps. But both failed in predictable ways when the reward function was hollow. Core The missing piece in Yang’s metaphor—and the reason it matters for crypto—is that RL only works if the reward function is well-designed. In AI, that means a dense, aligned signal that rewards the right behaviors and penalizes reward hacking. In crypto, the reward function is tokenomics. Every staking yield, every liquidity mining multiplier, every vesting schedule is a reward signal. Protocols that lean too far into RL (e.g., high-inflation yield farms) attract vampires—bots and mercenaries that game the system and dump. Protocols that over-index on SFT (e.g., heavy KYC, multisig bottlenecks, rigid governance) kill the explorative energy that drives innovation. The real insight from Yang’s framework is that most crypto teams operate in a muddle: they claim to be RL ("we empower contributors!") but their token incentives are SFT ("here's a fixed APY for locking"). The result is a misaligned environment where agents (both human and bot) learn to maximize the wrong metric. I have seen this firsthand in DAO grant committees: the reward function is “write a proposal that sounds ambitious,” so applicants learn to produce grandiose documents with zero delivery. That is reward hacking in the wild. Alchemy fails when the intent is hollow. What Yang glossed over—and what every crypto builder should internalize—is the problem of sparse rewards and credit assignment. In AI, long-horizon RL struggles because the agent gets a single reward at the end of a complex chain of actions. In crypto, that chain is a protocol’s development lifecycle. A team that rewards only the final product (TVL milestone, TGE, governance vote) creates a vacuum in the middle. Developers optimize for visible checkpoints, not for foundational infrastructure. The result: technical debt, security bugs, and ultimately a crash when the market turns bearish. The Lightning Network is a perfect example—seven years of RL exploration, but the reward function (routing reliability) was never properly aligned, so it remains a niche curiosity. Contrarian The contrarian take? Yang’s pro-RL stance is actually a luxury of a bull-market mindset. In a bear market, survival favors SFT. When liquidity is scarce and users are paranoid, you need guardrails: clear code standards, rigorous audits, predictable emission schedules. Pure RL without SFT constraints leads to chaos—ask Terra or FTX. The founder’s framing implicitly assumes a high-talent, high-trust environment (Moonshot’s current state), but most crypto projects are messy coalitions of anonymous anons, mercenary capital, and part-time contributors. For them, a heavy SFT layer isn’t oppression; it’s a lifeline. The best projects will be those that implement a “Constitutional” hybrid: a minimal set of immutable SFT rules (safety, audit requirements, core governance) and maximum RL freedom within those bounds. This is exactly what AI alignment researchers call Constitutional AI—and it is exactly what crypto needs to move beyond the current narrative of "decentralization or death." Takeaway The next narrative cycle will not be about which chain has the fastest finality or the most DAUs. It will be about which teams design the cleanest reward functions—both for their code and their people. Yang’s metaphor is a gift to crypto narrative hunters, but only if we strip it of its Silicon Valley gloss and apply it to the gritty reality of token incentives. The question every founder should ask: Is your protocol an RL agent exploring a rich environment, or a cat chasing a laser pointer you control? Bear markets burn through hollow intent. The alchemy that survives will be the one that aligns exploration with immutable constraints. Alchemy fails when the intent is hollow. Narratives are the only true non-fungible assets. A protocol without a story is just a smart contract waiting to die. — Narrative Hunter, Buenos Aires