Skip to main content

Skill memory in biped locomotion: Using perceptual information to predict task outcome – 1ª Parte

Page 1

Skill memory in biped locomotion Using perceptual information to predict task outcome J. Andre1, C. Santos2, L. Costa3 1 Department of Industrial Electronics, University of Minho, joaocandre@dei.uminho.pt 2 Department of Industrial Electronics, University of Minho, cristina@dei.uminho.pt 3 Department of Production and System, University of Minho, lac@dps.uminho.pt

Abstract Robots must be able to adapt their motor behavior to unexpected situations in order to safely move among humans. A necessary step is to be able to predict failures, which result in behavior abnormalities and may cause irrecoverable damage to the robot and its surroundings, i.e. humans. In this paper we build a predictive model of sensor traces that enables early failure detection by means of a skill memory. Specifically, we propose an architecture based on a biped locomotion solution with improved robustness due to sensory feedback, and extend the concept of Associative Skill Memories (ASM) to periodic movements by introducing several mechanisms into the training workflow, such as linear interpolation and regression into a Dynamical Motion Primitive (DMP) system such that representation becomes time invariant and easily parameterizable. The failure detection mechanism applies statistical tests to determine the optimal operating conditions. Both training and failure testing were conducted on a DARwIn-OP inside a simulation environment to assess and validate the failure detection system proposed. Results show that the system performance in terms of the compromise between sensitivity and specificity is similar with and without the proposed mechanism, while achieving a significant data size reduction due to the periodic approach taken. Keywords: Reinforcement learning · Bio-inspired · Skill Memory

robótica 99, 2.o Trimestre de 2015

UREËWLFD

$57,*2 &,(17õ),&2

1.ª Parte

1. INTRODUCTION Learning is a necessary skill for robots as it is with humans [21], and is still a demanding effort. It comprises an on-going process, and even after grasping the basic motions necessary for a specific movement, there is room for improvement either by trial-and-error or inference from perceptual data and previous experiences. Humans use the latter almost unconciously, constantly predicting their actions’ impact on their surroundings and improving them in order to achieve better outcomes. On robots, however, the use of sensor information is not nearly as automated, and it can be time-consuming to achieve proper parameterization with satisfactory results [20]. Robotic control is a challenging task due to the great number of variables at play in the interaction between a robot and its surroundings. Not being realistically possible to account in advance for all potential disruptive interference sources in the way movement execution, alternative ways to increase movement robustness must be found, either at hardware level (passive adaptation with compliant joints [22]) or at control level (extensive calibration, exploration though vision systems [22]).

Popular solutions are often based on a feedback loop that corrects movement plans as needed in response to perceptible changes in the environment [22]. Therefore an attractive enhancement in robotic systems would be the capability of using sensor information to improve its performance. Moreover, being able to identify potential failure conditions from perceptual information, while learning from past experiences, can prompt the robot to take countermeasures timely to ensure task completion, reacting to sudden fluctuations in the environment that otherwise would have a disruptive effect in movement execution. Several failure detection systems include particle [24] filters and neural-network based approaches [4, 5], however these are highly focused in industrial and wheeled robots, and up to th author’s knowledge there such a framework in humanoid/legged robots is lacking. This way it becomes necessary to find a way to store relevant perceptual information in consistently, in order to use it to monitor and improve robotic movement, what we would call a skill memory. Based on the premise that stereotypical movements tend to leave similar sensor footprints with each execution, even if the environment is dynamically changing, Pastor et al. have proposed the notion of Associative Skill Memory (ASM) [22]. An ASM saves a reference sensor footprint illustrating the most common values and variation of perceptual data when executing a certain task - associative because the memory itself is not solely based on sensor information, but rather on the association of this specific sensor footprint with the corresponding movement, skill or end task [21]. Overall, this work presents the first step towards a truly sensor-driven CPG biped locomotion approach, using ASMs to construct a vast library of locomotor movements and parameters, selectable according to the external context to achieve the most apropriate gait. There are however several problems to deal with, and failures are highly undesirable for autonomous humanoid robots expected to cope with realistic and possibly unforeseen physically interactive human environments. Throughout this work we assume failures as any deviation from optimal behavior. In the specific case of biped locomotion, this includes, not only falling but also unstable or inefficient walking patterns. During a robotic walk, these failures result in abnormalities in the robot’s normal behavior and may lead to falls and consequently irrecoverable damage to the robot and its surroundings, often populated by humans. In this paper we present a method to construct a predictive model for locomotion from sensor traces acquired during past attempts. This method is explored to reliably predict potential locomotion failures. This ASM-based failure predictor module


$57,*2 &,(17õ),&2 UREËWLFD

and prompted us to store the ASM on a period basis, evolving from a time-indexed reference over P motion periods to a phase-indexed data set representing a period of the sensor data variation (during a gait cycle), which can be loaded at the beginning of each period, eliminating the issues associated with movement periodicity. Additionally, this structure adds time-invariance to the locomotion task - the skill memory does not depend on the time but rather on the period or phase of the gait cycle. When considering such a data normalization process over the gait cycle, however, several remarks can be made: 1) period duration may not be constant during the complete movement, 2) movement may start in mid-period and 3) a movement evolution reference, such as an oscillator phase, is required. The data normalization procedure we propose, for the DARwIn-OP specifically replaces the time index associated with each sensor data value by the correspondent CPG right leg oscillator phase value (left leg oscillator was discarded as it begins in mid-cycle in the CPG system). Then linear interpolation is performed to achieve a regularly spaced waveform with Kphase datapoints representing a single period, posteriorly converted to percentage values for a more intuitive and generic representation. Note that this representation is invariant to movement duration and thus better suited for an ASM. While we use the CPG rhythmic generator as an evolution reference, any type of sensor/measure that illustrates movement completion can be used. However, due to the sensory feedback mechanism implemented, the phase provided by the CPG is simultaneously modulated according to equation (11) and thus reflects the entrainment with the environment and body physics. This provides additional robustness to the algorithm as the feedback loop also influences interpolated data. The value of Kphase, set beforehand, influences phase resolution of the normalized data. Throughout this paper we have set Kphase = 200. 5.4. Cost-based Weighting After having an interpolated and phase-indexed period for each trial, in order to arrive at a single reference period, each trial has a different contribution based on the previously described cost function. Each trial after interpolation y(t) is evaluated according to a cost threshold, Cs, that separates successful trials from unsuccessful trials. All failures are immediately discarded, and the contribution of each successful trial is inversely proportional to the cost, according to the cost-based weighting described in Algorithm 1. The two phase-indexed datasets that compose the bulk of the ASM data for each sensor: and are then obtained. 5.5. DMP Regression Dynamic Motion Primitives (DMPs) are used to codify the phase-indexed ASM, and , into a smaller set of parameters. DMPs are non-linear dynamical systems that use attractor dynamics to design complex parametrized trajectories, making them robust against external perturbations, autonomous and easily modulated, and adaptable to both discrete and rhythmic movements by means of point and limit cycle attractors [9, 8]. DMPs serve to compactly represent policies since by changing its few parameters they adapt the resulting attractor landscapes.

Algorithm 1. Cost-based weighting.

A non-linear term f in the form of a weighted sum of D Gaussian kernels adjusts the DMP system through a set of the D kernel weights:

(14) (15)

This structure provides stability, flexibility and robustness to policy representation, and is linearly parameterizable [9, 8]. Therefore it is possible to encode all the ASM data into a parameterized function, which holds several advantages to conventional pointwise data storage: 1) the available sensor data becomes continuous, since it is approximated by dynamical systems; 2) data is easily manipulated, with all the particular advantages inherent to DMP systems; and 3) the total amount of data is significantly reduced, one of the most significant advantages of parametric representation [9, 8]. Taking into account all these features of DMP regression of the sensor data, we adopted a regression process similar to the one described by Pastor et al. [20, 19], albeit with some specific changes. First, data size reduction is taken even further by setting DMP kernel centers linearly spaced across the phase space with constant kernel width H, which is dependent on the number of kernels (H = 2.5D). This way the necessary parameters to reconstruct the sensor signal are only the damping constants x, y and βy, the kernel weights θ ; Y_D×1 [9, 8] and the DMP initial values. In our particular implementation of DMP regression, we used the Path Integral Policy Improvement with Black Box Optimization (PI2-BB) algorithm, which has been developed in earlier works and has show reasonable results [1], with cost function CPI2 (rn,j ) = (ytarget(n) −rn,j )2 , D = 100 and K = 500 [1] to generate an approximation of the sensor curve shape, ytarget . rn,j stands for the instant n of trial j. PI2-BB iteratively samples and evaluates new sets of D kernel weights that minimize the squared distance from the DMP to the sensor data. Overall 2 × S DMPs were built (S = 32), for and . This DMP regression step can be considered facultative, as it is not mandatory in ASM creation, and the only objective is to find an alternative and more compact way to represent the ASM - the data before DMP regression and after DMP reconstruction is essentially the same, albeit with the differences highlighted before.


Turn static files into dynamic content formats.

Create a flipbook