Skip to content

Latest commit

 

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

2025-08-04.13-01-45.mp4

Functional Deep Dive

env

The updated environment transitions tracking attributes from discrete options to a continuous space:

  • the environment observation maps coordinates directly into a combined space (shape=(dim+1,)) that stacks normalized wall clearances and instantaneous vehicle velocity data.-
  • uses iterative step counts (sim_code()) to align execution tracking. It splits operations into discrete steps (70 loops for acceleration calculations, 20 loops for directional steering modifications).
  • maps inputs to continuous coordinates (spaces.Box(low=-1, high=1, shape=(2,))) managing proportional adjustments to acceleration/braking and left/right steering angles.
  • appends processing variables to storage files (training_data2.pkl) using an automated pickling loop (pickle.dump) to log raw parameters alongside rewards.

actor

The network updates the basic setup with a probabilistic policy using structural components:

  • processes tensors through three sequential nn.LazyLinear layers (1024 cells wide) paired with nn.Tanh activations. The final layer feeds into a NormalParamExtractor module to isolate tracking mean (loc) and scaling parameters (scale).
  • wraps network outputs using a ProbabilisticActor configuration. It uses a TanhNormal structural profile to clip unbounded outputs cleanly within the legal environment bounds.
  • runs a parallel multi-layer regression network (three 1024-node hidden layers) that compresses inputs into a 1-dimensional value assessment.

loop

  • the SyncDataCollector streams batches of 100 frames (frames_per_batch). For each batch, the Generalized Advantage Estimator (GAE) computes baseline advantage scores across 50 internal optimization steps (num_epochs).
  • trajectory data is flattened using .reshape(-1), transferred to the CPU, and cached inside a ReplayBuffer. The pipeline extracts distinct, non-overlapping sub-batches of 20 transitions using a SamplerWithoutReplacement tracking model.
  • minimizes structural tracking losses via ClipPPOLoss. It accumulates policy objectives, value network errors, and exploration metrics before tracking gradient steps
  • enforces policy safety margins by clipping backpropagation updates to an absolute ceiling threshold (max_grad_norm = 1.0). Learning rates decay across training batches via a cosine annealing routine (CosineAnnealingLR).
  • suspends exploration parameters every 10 collection steps using a set_exploration_type(ExplorationType.DETERMINISTIC) context manager. This samples a 100-frame validation rollout to monitor convergence baselines.
Hyperparameter Metric Assigned Value Architectural Context
num_cells 1024 Width of individual hidden layers in both policy and value networks
lr 0.0001 Base step size constraint applied to the Adam optimization loop
sub_batch_size 20 Number of trajectory transitions sampled during inner loop epoch updates
num_epochs 50 Training iterations executed per collected batch of environment data
max_grad_norm 1.0 Absolute threshold value handling gradient clipping bounds
clip_epsilon 0.1 Clipping parameter regulating policy ratio changes
entropy_eps 0.01 Coefficient scaling the policy entropy bonus to prevent premature convergence

About

a precursor to my forza driving RL and how i was pre-testing a simulated env

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages