A Note on Reinforcement Learning to Develop Self-defined Agents' Behavior
Matteo Morini, Pietro Terna
Abstract
The key point in this note is the self-development of simple behavior strategies, consistently with the bounded rationality hypothesis. Our artificial agents adopt learning techniques, mainly unsupervised, to achieve internal consistency in their behavior, with unexpected results. Those results can be considered mainly as the effects of the observer interpretation. The first technique in use has the name Cross Targets: to train the learning agent we use data crossed between the guesses about the action to be done and the guesses about the following results. An application of the CT "blind" strategy development is then presented: random walkers solve a node classification problem on a graph, after having learnt how to remain in homogeneous regions.
Create a lesson
Related papers
The Local-to-Global AD-k Conjecture is Resolved
Wei Chen
Improved Methods for k-core Community Search
Ian Chen, Haotian Yi, Arun Sharma et al.
Graphlets as structural fingerprints of complex networks
Anna Pidnebesna, David Hartman, Aneta Pokorna et al.
WCCS: Efficient Wedge Conductance Community Search over Large Temporal Bipartite Graphs (Full Paper)
Longlong Lin, Wei Chen, Pingpeng Yuan et al.
Inferring Temporal Dependencies from Social Time Series with the Cross-Correlogram
Bridget Smart, Renaud Lambiotte, Takaaki Aoki et al.
On the Expressive Power of Implicit Line-Graph Higher-Order Weisfeiler--Leman
Fan Yang