Gittins index
The Gittins index is a measure of the reward that can be achieved through a given stochastic process with certain properties, namely: the process has an ultimate termination state and evolves with an option, at each intermediate state, of terminating. Upon terminating at a given state, the reward achieved is the sum of the probabilistic expected rewards associated with every state from the actual terminating state to the ultimate terminal state, inclusive. The index is a real scalar.
Gittins index - Wikipedia Jump to content From Wikipedia, the free encyclopedia Measure in decision theory The Gittins index is a measure of the reward that can be achieved through a given stochastic process with certain properties, namely: the process has an ultimate termination state and evolves with an option, at each intermediate state, of terminating. Upon terminating at a given state, the reward achieved is the sum of the probabilistic expected rewards associated with every state from the actual terminating state to the ultimate terminal state, inclusive. The index is a real scalar . Ter
Explore this link on the map →related reading
- Multi-armed bandit - Wikipediaen.wikipedia.org
- The Multi-Armed Bandit Problem and Its Solutions | Lil'Loglilianweng.github.io
- Optimal stopping - Wikipediaen.wikipedia.org
- Gregory Gundersengregorygundersen.com
- The Bayes Banditfrancesco215.github.io
- A (Long) Peek into Reinforcement Learning | Lil'Loglilianweng.github.io
- Stochastic process - Wikipediaen.wikipedia.org
- Pareto front - Wikipediaen.wikipedia.org
- RUDDER - Reinforcement Learning with Delayed Rewards | rudderml-jku.github.io
- Brian Christian on computer science algorithms that tackle fundamental and universal problems — and whether they can help us live better in practice | 80,000 Hours80000hours.org
- De Finetti's theorem - Wikipediaen.wikipedia.org
- On The Independence Axiom — LessWronglesswrong.com