The subproblem is to maximize the expected sum of rewards received while following an option plus a reward-weighted contribution of the feature value at termination, minus any penalty from bad terminal states, creating a trade-off between feature attainment and reward preservation.

definitionpending

Speaker

Richard Sutton

Evidence Quote

it's the sum... the sum is conditional on starting

Source

Rich Sutton, The OaK Architecture: A Vision of SuperIntelligence from Experience - RLC 2025Amii
Created: 8/11/2026, 7:21:57 AM

My Notes

Loading notes...