PyroGuard AI

Autonomous wildfire suppression using reinforcement learning and multi-drone coordination

What We Built

A complete wildfire simulation system with AI-powered drone agents

System Overview

1

Wildfire Environment

2D visualization with realistic fire spread, wind effects, and terrain

2

Rule-Based Drones

Autonomous patrol, fire detection, and suppression with resource management

3

RL-Trained Drones

Deep Q-Network agents with dueling architecture and experience replay

4

Multi-Drone Coordination

Fleet management with shared intelligence and optimized resource allocation

Key Metrics

30x30
Grid Size
3-6
Drones
DQN
AI Model

Implemented Features

Core functionality we actually built

Fire Dynamics

Heat-based spread with wind influence, multiple intensity levels, and terrain-dependent burn rates

Autonomous Drones

Patrol patterns, fire detection, water management, and emergency recharging behavior

RL Training

Dueling DQN with CNN processing, prioritized experience replay, and multi-objective rewards

Multi-Drone System

Coordinated fleet operations with different drone types (simple + RL) working together

Performance Analytics

Comprehensive metrics logging, training visualization, and real-time statistics

Interactive Controls

Manual drone control, fire spawning, wind changes, and real-time parameter adjustment

Training & Reward System

How we trained the AI with sophisticated reward engineering

Reward Structure

Fire Suppressed +10.0

Primary objective - extinguishing fires

Tree Saved +2.0

Protecting vegetation from fire spread

Episode Success +100.0

Complete mission success bonus

Proximity to Fire +0.2

Encourage movement towards fires

Coverage Bonus +0.5

Encourage exploration of new areas

Penalty System

Crash Penalty -50.0

Moving out of bounds or dangerous areas

Time Penalty -0.1

Encourage efficient time usage

Fire Spread -2.0

Penalty for each new fire that appears

Vegetation Loss -0.5

Penalty for each tree/vegetation burned

Training Process

Episodes: 3,000-5,000
Max Steps: 400 per episode
Learning Rate: 3e-4
Gamma: 0.99

Network Architecture

Type: Dueling DQN
CNN Layers: 3 layers
Hidden Size: 512 neurons
Replay Buffer: 50,000 experiences

Training Results

Success Rate: 85%+
Avg Reward: 150+
Training Time: 2-4 hours
Convergence: ~2000 episodes

Performance Comparison

Single rule-based vs Multiple RL agents in similar settings

Single Rule-Based Agent

Traditional approach with fixed logic

Fire Detection 5 cell radius
Decision Making Fixed rules
Coordination None
Adaptability Limited
Success Rate ~45%
Avg Fires Extinguished 3.2 per episode

Multiple RL Agents

AI-powered coordinated fleet

Fire Detection 5 cell radius
Decision Making Deep Q-Network
Coordination Shared intelligence
Adaptability High
Success Rate ~85%
Avg Fires Extinguished 7.8 per episode
89%
Improvement in Success Rate
RL agents vs rule-based
2.4x
More Fires Extinguished
Per episode average
3-6
Coordinated Drones
Working together

Key Advantages of RL Approach

1

Adaptive Learning

Agents learn from experience and improve over time

2

Multi-Objective Optimization

Balances fire suppression, efficiency, and safety

3

Coordinated Strategy

Drones work together with shared intelligence

4

Resource Efficiency

Optimized water and energy usage patterns

5

Dynamic Response

Adapts to changing fire conditions and wind

6

Scalable Performance

More drones = better coverage and efficiency

Technology Stack

Technologies we actually used

🐍

Python 3.11

Core language

🔥

PyTorch

Deep learning

🎮

Gymnasium

RL environment

🎨

Pygame

3D visualization

📊

Matplotlib

Training plots

🧮

NumPy

Numerical computing

🤖

DQN

RL algorithm

👁️

CNN

Spatial processing