This guide assumes you are familiar with ap_gym. If you are not, please refer to the ap_gym documentation.
In tactile perception environments, the agent has to identify properties of 3D objects by exploring them with a GelSight Mini tactile sensor. The agent does not have access to the location of the objects and also receives no visual input. Instead, it must actively control the sensor to find and explore them.
Currently implemented are the following tasks, which are described in more detail in their respective documentations:
![]() TactileMNIST-v0 |
![]() Starstruck-v0 |
![]() Toolbox-v0 |
![]() TactileMNISTVolume-v0 |
For an example usage of tactile perception environments, see example/tactile_mnist_env.py.
All tactile perception environments are instantiations of the tactile_mnist.TactilePerceptionVectorEnv class and share the following properties:
| Action Space |
Dict({ "sensor_target_pos_rel":Box(-1.0, 1.0, shape=(3,), dtype=np.float32), "sensor_target_rot_rel": Box(-1.0, 1.0, shape=(6,), dtype=np.float32)})"sensor_target_rot_rel" is only present if the environment allows sensor orientation.
|
| Observation Space |
Dict({ "sensor_img": ap_gym.ImageSpace(width=W, height=H, channels=3, dtype=np.float32), "sensor_pos": Box(-1.0, 1.0, shape=(3,), dtype=np.float32) "sensor_rot": Box(-1.0, 1.0, shape=(6,), dtype=np.float32) "time_step": Box(-1.0, 1.0, shape=(), dtype=np.float32)})"sensor_rot" is only present if the environment allows sensor orientation.
|
The action is a dictionary with the following keys:
| Key | Type | Description |
|---|---|---|
"sensor_target_pos_rel" |
np.ndarray |
3-element float32 numpy vector containing the normalized relative linear target movement of the sensor in the range |
"sensor_target_rot_rel" |
np.ndarray |
3-element float32 numpy vector containing the normalized relative rotational target movement of the sensor as rotation vector in the range |
To compute the next sensor position, the "sensor_target_pos_rel" value is first projected into the unit sphere and multiplied with the maximum distance the sensor can move in one step.
The maximum sensor movement is computed from the transfer_timedelta_s, linear_acceleration, and linear_velocity parameters.
The resulting vector is then added to the current sensor position.
To ensure that the sensor is in contact with the object while not penetrating it, the sensor is moved towards or away from the target position until it touches the object or the cell surface.
This movement is always perpendicular to the sensor's sensing surface.
To compute the next sensor orientation, the "sensor_target_rot_rel" value is first projected into the unit sphere and multiplied with the maximum rotation the sensor can perform in one step.
The resulting rotation is then multiplied with the current sensor orientation to get the new target sensor orientation.
Since we only want to allow to point downwards, we ensure that the angle between the sensor's z-axis and the world's z-axis is at maximum max_tilt_angle (default: 90 degrees).
This constraint is enforced by projecting the target sensor orientation back into the valid orientation space if it exceeds the maximum tilt angle, which yields the final sensor orientation.
The observation is a dictionary with the following keys:
| Key | Type | Description |
|---|---|---|
"sensor_img" |
np.ndarray |
float32 representing a tactile reading where each pixel is in the range |
"sensor_pos" |
np.ndarray |
3-element float32 numpy vector containing the normalized position of the sensor in the range |
"sensor_rot" |
np.ndarray |
6-element float32 numpy vector containing the orientation of the sensor in the range |
"time_step" |
float |
The current time step between 0 and step_limit normalized to the range "terminate" (default). |
We model 3D orientations as 6D vectors as suggested by Zhou et al. (2019). Unlike this work though, we include the second and third column of the rotation matrix instead of the first and second, as it helps us to ensure that the sensor only receives downwards pointing target orientations.
The reward at each timestep is a sum of:
- A small action regularization equal to
$10^{-3} \cdot{} \lVert action\rVert$ . - The loss of the current prediction of the agent
The tactile sensor starts at a randomly sampled pose in the workspace.
Specifically, the position is uniformly randomly samples from
The episode ends with the terminate flag set when the maximum number of steps (step_limit) is reached.
Here is an example of how to use the environments:
import ap_gym
import numpy as np
env = ap_gym.make("tactile_mnist:TactileMNIST-v0")
# Alternatively:
# env = ap_gym.make("tactile_mnist:Starstruck-v0")
env.reset()
for i in range(10):
action = {
"action": {
"sensor_target_pos_rel": np.random.uniform(-1, 1, size=3),
# Uncomment the following line if the environment uses sensor orientation
# "sensor_target_rot_rel": np.random.uniform(-1, 1, size=6),
},
"prediction": np.ones(10)
}
obs, _, _, _, info = env.step(action)
sensor_img = obs["sensor_img"] # 64 x 64 x 3 tactile image
sensor_pos = obs["sensor_pos"] # Normalized 3D sensor position
time_step = info["time_step"] # Normalized current time step
# Only if the environment uses sensor orientation
# sensor_rot = obs["sensor_rot"] # 6D sensor orientation
ground_truth_label = info["prediction"]["target"] # integer ground truth label of the objectA full example can be found in example/tactile_mnist_env.py or example/tactile_mnist_env_vec.py for the vectorized version.
| Parameter | Type | Default | Description |
|---|---|---|---|
config |
TactilePerceptionConfig |
Configuration of the tactile perception environment. See the TactilePerceptionConfig documentation for details. | |
num_envs |
int |
Number of parallel environments to create. | |
single_prediction_space |
gym.Space[PredType] |
The prediction space of the environment. | |
single_prediction_target_space |
gym.Space[PredTargetType] |
The prediction target space of the environment. | |
loss_fn |
ap_gym.LossFn |
The loss function of the environment. | |
render_mode |
Literal["rgb_array", "human"] |
"rgb_array" |
Which render mode to use. |
There are currently two types of tactile perception environments: Tactile Classification Environments and Tactile Regression Environments. Check out their respective documentations for more details.



