coach

gryf/coach

mirror of https://github.com/gryf/coach.git synced 2026-07-07 01:46:31 +02:00

Author	SHA1	Message	Date
Itai Caspi	a7206ed702	Multiple improvements and bug fixes (#66 ) * Multiple improvements and bug fixes: * Using lazy stacking to save on memory when using a replay buffer * Remove step counting for evaluation episodes * Reset game between heatup and training * Major bug fixes in NEC (is reproducing the paper results for pong now) * Image input rescaling to 0-1 is now optional * Change the terminal title to be the experiment name * Observation cropping for atari is now optional * Added random number of noop actions for gym to match the dqn paper * Fixed a bug where the evaluation episodes won't start with the max possible ale lives * Added a script for plotting the results of an experiment over all the atari games	2018-02-26 12:29:07 +02:00
Zach Dwiel	ef46e194af	remove unused commented code	2018-02-21 10:05:57 -05:00
Zach Dwiel	d9303e731e	remove python2 compatibility	2018-02-21 10:05:57 -05:00
Zach Dwiel	5cf10e5f52	fix bug in ddpg	2018-02-21 10:05:57 -05:00
Zach Dwiel	8248caf35e	fix more agents	2018-02-21 10:05:57 -05:00
Zach Dwiel	98f57a0d87	fix ddpg	2018-02-21 10:05:57 -05:00
Zach Dwiel	ee6e0bdc3b	fix keep_dims -> keepdims	2018-02-21 10:05:57 -05:00
Zach Dwiel	39a28aba95	fix clipped ppo	2018-02-21 10:05:57 -05:00
Zach Dwiel	85afb86893	temp commit	2018-02-21 10:05:57 -05:00
Gal Leibovich	16c5032735	fix for tensorboard visualization slowing execution even when it is off apparently tensorflow still collect summary data even when no summary FileWriter is defined.	2018-02-18 16:35:24 +02:00
Gal Leibovich	7c8962c991	adding support in tensorboard (#52 ) * bug-fix in architecture.py where additional fetches would acquire more entries than it should * change in run_test to allow ignoring some test(s)	2018-02-05 15:21:49 +02:00
Itai Caspi	43821c9630	adding the selu activation	2018-01-22 12:05:43 +02:00
Zach Dwiel	6c79a442f2	update nec and value optimization agents to work with recurrent middleware	2018-01-05 20:16:51 -05:00
Itai Caspi	125c7ee38d	Release 0.9 Main changes are detailed below: New features - * CARLA 0.7 simulator integration * Human control of the game play * Recording of human game play and storing / loading the replay buffer * Behavioral cloning agent and presets * Golden tests for several presets * Selecting between deep / shallow image embedders * Rendering through pygame (with some boost in performance) API changes - * Improved environment wrapper API * Added an evaluate flag to allow convenient evaluation of existing checkpoints * Improve frameskip definition in Gym Bug fixes - * Fixed loading of checkpoints for agents with more than one network * Fixed the N Step Q learning agent python3 compatibility	2017-12-19 19:27:16 +02:00
Itai Caspi	11faf19649	QR-DQN bug fix and imporvements (#30 ) * bug fix - QR-DQN using error instead of abs-error in the quantile huber loss * improvement - QR-DQN sorting the quantile only once instead of batch_size times * new feature - adding the Breakout QRDQN preset (verified to achieve good results)	2017-11-29 14:01:59 +02:00
Zach Dwiel	9ae2905a76	clean up input embeddings setup	2017-11-14 17:39:18 +02:00
galleibo-intel	3c330768f0	Fix for NEC not saving the DND when saving a model	2017-11-09 19:13:23 +02:00
cxx	84e536d371	Fix std calculation using unbiased estimation in sharing stat mode.	2017-11-07 20:19:54 +02:00
Itai Caspi	b40259c61a	bug fix - remove import warning when everything was imported successfully + changed global step api to match TF 1.4	2017-11-06 17:28:13 +02:00
Itai Caspi	a8bce9828c	new feature - implementation of Quantile Regression DQN (https://arxiv.org/pdf/1710.10044v1.pdf ) API change - Distributional DQN renamed to Categorical DQN	2017-11-01 15:09:07 +02:00
Itai Caspi	913ab75e8a	bug fix - preventing crashes when the probability of one of the actions is 0 in the policy head	2017-10-31 10:51:48 +02:00
Itai Caspi	1918f16079	imporved API for getting / setting variables within the graph	2017-10-31 10:51:48 +02:00
cxx	f43c951c2d	Unify base class using new-style (object).	2017-10-26 12:33:09 +03:00
Gal Leibovich	1d4c3455e7	coach v0.8.0	2017-10-19 13:10:15 +03:00

24 Commits