Dynamic Control of Synchronised Automata using Networked Q-Learning Techniques
C. Deeks
https://doi.org/10.19124/ima.2015.001.22
Abstract
This paper describes a novel approach to the control of a number of autonomous robots operating concurrently but without direct communication or coordination between them. Each automaton is a part of each of the others’ dynamic local environments, and control actions need to be selected in response to observations of this local environment. A variation on Qlearning is introduced where the set of actions available to an automaton can be artificially constrained in response to observations and so with learned policies being recalculated online. The model to which this approach has been applied is a representation of a set of production line automata, where there may be a desire for a greater degree of concurrent operation to make the whole production line more efficient, but where there is also the desire to avoid long ramp-up or configuration times. It is shown that robust behaviour within set geometric constraints can be observed, even when the problem is set so that the automata cannot avoid obstructing each other. Training not just one automaton with a machine learning technique, but training a family of them concurrently, is therefore shown to be a feasible candidate for more widespread application in coordinated control tasks in dynamic environments.
