gryf/coach

mirror of https://github.com/gryf/coach.git synced 2026-03-18 15:53:35 +01:00

Files

Gal Leibovich 6e08c55ad5 Enabling-more-agents-for-Batch-RL-and-cleanup (#258 )

allowing for the last training batch drawn to be smaller than batch_size + adding support for more agents in BatchRL by adding softmax with temperature to the corresponding heads + adding a CartPole_QR_DQN preset with a golden test + cleanups

2019-03-21 16:10:29 +02:00

__init__.py

pre-release 0.10.0

2018-08-13 17:11:34 +03:00

basic_rl_graph_manager.py

Batch RL (#238 )

2019-03-19 18:07:09 +02:00

batch_rl_graph_manager.py

Enabling-more-agents-for-Batch-RL-and-cleanup (#258 )

2019-03-21 16:10:29 +02:00

graph_manager.py

Batch RL (#238 )

2019-03-19 18:07:09 +02:00

hac_graph_manager.py

removing datasets + imports optimization

2018-08-27 10:54:11 +03:00

hrl_graph_manager.py

removing datasets + imports optimization

2018-08-27 10:54:11 +03:00

README.md

pre-release 0.10.0

2018-08-13 17:11:34 +03:00

README.md

Block Factory

The block factory is a class which creates a block that fits into a specific RL scheme. Example RL schemes are: self play, multi agent, HRL, basic RL, etc. The block factory should create all the components of the block and return the block scheduler. The block factory will then be used to create different combinations of components. For example, an HRL factory can be later instantiated with:

env = Atari Breakout
master (top hierarchy level) agent = DDPG
slave (bottom hierarchy level) agent = DQN

A custom block factory implementation should look as follows:

class CustomFactory(BlockFactory):
    def __init__(self, custom_params):
        super().__init__()

    def _create_block(self, task_index: int, device=None) -> BlockScheduler:
        """
        Create all the block modules and the block scheduler
        :param task_index: the index of the process on which the worker will be run
        :return: the initialized block scheduler
        """

        # Create env
        # Create composite agents
        # Create level managers
        # Create block scheduler

        return block_scheduler