Exploring Raspberry Pi Pico PIO
In early 2021, the Raspberry Pi Foundation launched the Raspberry Pi Pico microcontroller. At the time, it felt novel because Raspberry Pi hadn’t had a product line like this before. I bought one, but ended up shelving it without digging into it much.
Upon closer inspection, I realized the specs were actually pretty impressive:
| Raspberry Pi pico | Specs | Arduino nano |
|---|---|---|
| ARM M0+ dual core | MCU | ATmega328p |
| 133MHz | Clock Speed | 16MHz |
| 32-bit | Architecture | 8-bit |
| 264kB | SRAM | 2.5kB |
| 2MB (RP2040 本身無 flash) | Flash | 32kB |
| UARTx2 USB host 1.1 SPI Timer RTC | Peripherals | UARTx1 SPI Timer |
With dual cores, 32-bit architecture, 264kB of RAM, and 2MB of flash memory, it blows the Arduino Nano out of the water and even outperforms some STM32 chips. It’s also remarkably cheap—accessible for just over 100 TWD (around $4 USD). Beyond C/C++, it supports MicroPython and even offers debugging capabilities, allowing you to set breakpoints and inspect memory or variable states via GDB.
Another feature that amazed me is PIO (Programmable I/O). But before diving into that, let’s talk about GPIO and the problem PIO is trying to solve!
What is GPIO?
GPIO stands for General Purpose Input/Output. In microcontrollers, it typically provides the ability to control pin output or read pin input, letting you programmatically set a pin high or low.
A simple example is blinking an LED. If you want to make an LED blink, you can connect one lead to ground and the other to a GPIO pin, then program the pin to alternate between high and low voltage states. Beyond toggling LEDs, GPIO pins are also used for data communication, such as I2C or UART.
Peripherals
To allow microcontrollers to communicate with external peripherals, they usually come with built-in hardware implementations of common communication protocols. For instance, the Arduino supports UART; if you use a Pro Micro, its onboard AVR ATmega32U4 chip even has built-in USB support ready to use.
The downside, however, is that if a microcontroller doesn’t natively include a specific protocol peripheral, developers have to either buy dedicated ICs or implement the protocol manually via GPIO pins. This concept is somewhat akin to the difference between hardware decoding and software decoding.
For example, on Arduino, we can use SoftwareSerial to implement UART at the software level. I used this in my post last year on building an Arduino CO2 sensor1(/devnote/2020-07-24-arduino-esp32-co2-sensor-2/):
// https://github.com/kjj6198/MH-Z14A-arduino/blob/master/co2.ino#L14
...
SoftwareSerial co2Serial(3, 4); // RX, TX
co2Serial.write(commands, 9); // send command
co2Serial.readBytes(response, 9);
Under the hood, SoftwareSerial uses GPIO pins to emulate the UART protocol. The benefit is that the Arduino’s native hardware UART can remain connected to the computer for debugging, while SoftwareSerial handles communication with other external devices.

Data Transmission Relies on Precise Timing Control
Hardware data transmission relies heavily on timing—often needing to be accurate down to CPU clock cycles to prevent transmission errors. In the implementation of SoftwareSerial:
void SoftwareSerial::begin(long speed)
{
// 略
// Precalculate the various delays, in number of 4-cycle delays
uint16_t bit_delay = (F_CPU / speed) / 4;
// 12 (gcc 4.8.2) or 13 (gcc 4.3.2) cycles from start bit to first bit,
// 15 (gcc 4.8.2) or 16 (gcc 4.3.2) cycles between bits,
// 12 (gcc 4.8.2) or 14 (gcc 4.3.2) cycles from last bit to stop bit
// These are all close enough to just use 15 cycles, since the inter-bit
// timings are the most critical (deviations stack 8 times)
_tx_delay = subtract_cap(bit_delay, 15 / 4);
// Only setup rx when we have a valid PCINT for this pin
if (digitalPinToPCICR((int8_t)_receivePin)) {
#if GCC_VERSION > 40800
// Timings counted from gcc 4.8.2 output. This works up to 115200 on
// 16Mhz and 57600 on 8Mhz.
//
// When the start bit occurs, there are 3 or 4 cycles before the
// interrupt flag is set, 4 cycles before the PC is set to the right
// interrupt vector address and the old PC is pushed on the stack,
// and then 75 cycles of instructions (including the RJMP in the
// ISR vector table) until the first delay. After the delay, there
// are 17 more cycles until the pin value is read (excluding the
// delay in the loop).
// We want to have a total delay of 1.5 bit time. Inside the loop,
// we already wait for 1 bit time - 23 cycles, so here we wait for
// 0.5 bit time - (71 + 18 - 22) cycles.
_rx_delay_centering = subtract_cap(bit_delay / 2, (4 + 4 + 75 + 17 - 23) / 4);
// There are 23 cycles in each loop iteration (excluding the delay)
_rx_delay_intrabit = subtract_cap(bit_delay, 23 / 4);
// There are 37 cycles from the last bit read to the start of
// stopbit delay and 11 cycles from the delay until the interrupt
// mask is enabled again (which _must_ happen during the stopbit).
// This delay aims at 3/4 of a bit time, meaning the end of the
// delay will be at 1/4th of the stopbit. This allows some extra
// time for ISR cleanup, which makes 115200 baud at 16Mhz work more
// reliably
_rx_delay_stopbit = subtract_cap(bit_delay * 3 / 4, (37 + 11) / 4);
#else // Timings counted from gcc 4.3.2 output
// Note that this code is a _lot_ slower, mostly due to bad register
// allocation choices of gcc. This works up to 57600 on 16Mhz and
// 38400 on 8Mhz.
_rx_delay_centering = subtract_cap(bit_delay / 2, (4 + 4 + 97 + 29 - 11) / 4);
_rx_delay_intrabit = subtract_cap(bit_delay, 11 / 4);
_rx_delay_stopbit = subtract_cap(bit_delay * 3 / 4, (44 + 17) / 4);
#endif
...
tunedDelay(_tx_delay); // if we were low this establishes the end
}
...
}
The code isn’t very long, but to hit the right timing, it even accounts for the CPU cycles spent by different GCC versions, subtracting them before executing the delay. This illustrates just how critical timing is for data transmission. While you could implement this using hardware timers and interrupts instead, hardware timers are also limited in number.
Bit Banging
Being able to implement communication protocols in pure software is convenient, but the downside is that it is very resource-intensive for the processor. The higher the communication frequency, the more CPU cycles are burned managing cycle-accurate timing. If you need precise timing outputs or want to keep your CPU free from protocol overhead, PIO is designed to solve exactly this problem.
PIO (Programmable I/O)
Overview
As mentioned earlier, the problem is that protocol timing demands consume precious processor resources. PIO can satisfy these timing requirements up to the processor’s full clock frequency (133MHz) without burdening the main CPU. You can think of PIO as small coprocessors dedicated specifically to GPIO pins. These small processors operate independently without stealing main CPU cycles, communicating with the main processor via FIFOs and IRQs.
Each RP2040 contains two PIO blocks, and each block houses 4 state machines. Every state machine can be reconfigured dynamically through code to implement different communication interfaces.
PIO provides a simplified assembly language with just 9 instructions, two registers, and a maximum program length of 32 instructions. Despite its compactness, this instruction set is capable of satisfying most communication protocol requirements.

(Image source: RP2040 Datasheet)
As you can see from the diagram, the four state machines share the same program memory. Because the instruction memory has four read ports, each state machine can access instructions simultaneously without blocking one another.
Introducing the State Machine
Each PIO block contains four state machines that share the same program memory. However, each state machine can be configured to target different GPIO pins. For example, if you implement a UART program, the 4 state machines allow you to instantiate up to four completely independent UART interfaces.
A State Machine consists of the following components:
- OSR (Output Shift Register): 32-bit; receives data from the main processor via the TX FIFO.
- ISR (Input Shift Register): 32-bit; transfers data to the main processor via the RX FIFO.
- X and Y registers: Two general-purpose registers per state machine.
- PC: Program counter.
- Clock divider: State machines can run up to the main processor’s clock speed, which is too fast for most protocols. The clock divider lets you scale the frequency down (ranging from 1 to 65536).
- Instructions

IO Mapping
IO mapping is slightly more intricate than on other microcontrollers. It might feel a bit convoluted at first, but once you grasp it, the design makes complete sense. Each IO operation falls into one of four roles: input, output, set, and sideset.
- input: Reads data from external sensors or devices (similar to
digitalReadin Arduino). - output: Controls pin voltage levels via software (similar to
digitalWritein Arduino). - set: Sets pin output levels directly (similar to output, but with subtle differences).
- sideset: Alters the state or direction of other pins concurrently while executing an instruction.
Set and sideset might be the trickiest to understand, and we’ll dive deeper into them shortly. A single GPIO pin can belong to multiple groups simultaneously; for example, you can assign a GPIO to be both an input and an output.
Each IO mapping is configured using a base pin and a pin count. For example, to configure GPIO0 and GPIO1 as SET pins, set the base pin to GPIO0 and the count to 2. This implies that pins within any given role mapping must be contiguous; you cannot, for instance, define an OUTPUT group consisting of non-contiguous pins like GPIO0, GPIO3, and GPIO5.
INPUT and OUTPUT support up to 32 pins (even though the Pico board only exposes 30 GPIOs). SET and SIDESET support up to 5 pins.

In summary, IO mapping has several key characteristics:
- A single pin can belong to multiple roles at the same time (e.g., both SET and OUTPUT).
- INPUT and OUTPUT support up to 32 pins, while SET and SIDESET support up to 5 pins.
- Pin groups must be contiguous (e.g., from GPIO0 to GPIO3).
IRQ (Interrupt Request)
IRQ flags can be used to trigger interrupts or synchronize state across different state machines.
Introducing PIO Assembly Language
PIO provides a simple yet powerful assembly language with only 9 instructions:
- SET
- IN
- OUT
- PULL
- PUSH
- JMP
- WAIT
- MOV
- IRQ
The coding style is essentially identical to standard assembly, so I won’t linger on basic syntax. However, there are a few operands/registers in PIO assembly worth remembering upfront:
- pins: Represents the pins mapped to this PIO state machine. For instance, if mapped starting at GPIO0, pin 0 is GPIO0; if mapped starting at GPIO2, pin 0 is GPIO2.
- pindirs: Configures pin direction.
0for input,1for output. - X, Y: General-purpose registers.
- osr: Output Shift Register.
- isr: Input Shift Register.
- data: Immediate data, up to 5 bits (i.e., values 0 to 31).
Having registers and jump instructions fulfills the basic requirements for Turing completeness. In theory, you could implement arithmetic operations with PIO, though PIO wasn’t designed for general computation—it’s more of a fun thought experiment.
Delay Functionality
To achieve precise timing control without wasting precious instruction memory, PIO provides a very handy feature: appending [] after an instruction to specify the number of delay cycles. For example:
loop:
set pins, 1 [1] ; set 需要 1cycle,除此之外再等 1cycle
set pins, 0
jmp loop
This is equivalent to:
loop:
set pins, 1
nop
set pins, 0
jmp loop
This lets us hold pins high for 2 cycles and low for 2 cycles without stuffing nop instructions and wasting instruction space. The value inside [] can range from 1 to 31.
Side-Set
Side-set allows you to modify the voltage levels of side-set pins concurrently with the execution of an instruction. You must explicitly declare the number of side-set pins in your program:
.side_set 1
This indicates that 1 pin will be used for side-set. You can specify up to 5 side-set pins, which can overlap with other mappings (such as input or output).
This is particularly convenient during state transitions. For example, UART stays high when idle, while the start bit is low. We can change the pin state at the exact moment we pull data (example adapted from pico-examples/uart_tx.pio):
.program uart_tx
.side_set 1
pull side 1 [7]
set x, 7 side 0 [7]
loop:
out pins, 1
jmp x-- bitloop [6]
In this example, every time pull is executed, the side-set pin is set to 1 simultaneously, serving as the stop bit. When initializing the X register, side-set can set the pin to 0 directly as the start bit. This eliminates the need to spend an extra instruction and cycle on pin toggling, which is very convenient.
Here, 8 cycles are allocated per bit. Within each bit period, up to 8 instructions can be executed; any unused cycles can easily be padded out using delays to achieve exact timing.
Note that using side-set reduces the number of bits available for delay cycles in instruction encodings. For instance:
.side_set 3 ; 將兩個腳位設定為 side_set
set x, 1 side 0 [3] ; 因為 3bit 已經拿去給 side_set,delay 最多 2bit(5-3),也就是 1~3
SET

Writes data to a destination. If the destination is pins, it targets the pins mapped to SET.
set X, 30 ; 設定 X 為 30
set pins, 1 ; 設定 SET 腳位為 1
set pins, 5 ; 設定 SET 腳位為 5(以二進位來說會是 0b101)
IN

Shifts a specified number of bits from source into the ISR. For example:
in osr, 1
Shifts 1 bit out of the OSR into the ISR. Or:
in pins, 4
Reads 4 bits from the input pins (starting from the base pin) and shifts them into the ISR.
OUT

Shifts data out of the OSR into the destination. For example:
out pins, 2
Takes 2 bits from the OSR and writes them to the output pins.
PULL

Reads 32 bits from the TX FIFO into the OSR. If no parameters are provided to pull, by default it will wait until data arrives in the TX FIFO (it does not need to fill all 32 bits) before continuing execution. For example:
loop:
pull ; 等到 tx 有資料後才會繼續執行,沒有的話會一直停在這行
out pins ; 將 OSR 的資料放到 output pins
jmp loop
Additional modifiers can be specified:
ifempty: Executes only when the OSR is empty; otherwise does nothing.block: Stalls and waits if the TX FIFO is empty.noblock: Copies the contents of register X into the OSR if the TX FIFO is empty (equivalent toMOV OSR, X).
Protocols like UART wait continuously when no data is being transmitted and only begin processing once data arrives; using block easily achieves this behavior.
PUSH

Pushes the contents of the ISR into the RX FIFO and clears the ISR to zero.
iffull: Executes only when the ISR is full; otherwise does nothing.block: Stalls if the RX FIFO is full.
JMP

Jumps to a target address if the condition is met. PIO’s jmp supports several conditions:
- No condition: When no condition is specified, it executes as an unconditional jump (
always). !X: Jump ifX == 0.X--: Jump ifX != 0, decrementingXon each execution.!Y: Jump ifY == 0.Y--: Jump ifY != 0, decrementingYon each execution.X!=Y: Jump ifX != Y.PIN: Jump based on the input pin’s logic level (configured via thesm_config_set_jmp_pinfunction):- Jumps if high.
- Does not jump if low.
!OSRE: Jump if OSR is not empty.
loop:
set x, 30 ; 設定 x
jmp x-- loop; 當 x 不為 0 時跳到 loop,同時 x-1
WAIT

Stalls until the specified condition becomes true. Key parameters include:
- GPIO: Selects an absolute GPIO pin by index, independent of the state machine’s IO mapping.
- PIN: Selects an input pin by index, relative to the state machine’s IO mapping.
- IRQ: Waits until the IRQ flag at the specified index matches the polarity before executing the next instruction.
wait 1 pin 0 ; 等待 input pin0 為 1 後才執行下一行程式碼
wait 0 gpio 1 ; 等待 GPIO1 為 0 後才執行下一行程式碼
MOV

Copies data from a source to a destination. Two special operators are available to assist:
!or~: Performs a bitwise NOT during copy.::: Reverses the bit order (bit-reverse) during copy.
This design also helps conserve valuable instruction memory.
mov X, Y ; 將 Y 複製到 X
mov pins, X ; 將 X 複製到 pins
mov pins, ::X ; 將 X 的 bit 反轉後複製到 pins
IRQ

Sets or clears an IRQ flag. Once an IRQ flag is set, it can trigger an interrupt depending on how the main program is configured.
irq wait 1 rel
Available parameters include:
- wait: Waits until the flag is cleared before proceeding.
- nowait: Continues execution immediately without waiting for the flag to clear.
- clear: Clears the specified IRQ flag.
If no parameter is specified, the default is nowait. Adding rel adds the current state machine index (modulo 4) to the IRQ index, allowing the main program to handle interrupts on a per-state-machine basis.
Specifying Program Execution Range
Unless stopped, a state machine will wrap back to the beginning after reaching the end of the program. However, we can use specific directives to instruct the PIO where execution should wrap:
set x, 8
.wrap_target
set pins, 1
set pins, 0
.wrap
Wrapping instructions with .wrap_target and .wrap defines the loop boundary for the state machine, saving an instruction that would otherwise be spent on jmp.
Integrating PIO with Main Application Code (C/C++)
The pico-examples repository provides numerous PIO examples. Here, let’s look at a simple blink implementation.2
First, we write the PIO code in a file with the .pio extension:
.program blink
.wrap_target
set pins, 1
nop [19]
nop [19]
set pins, 0
nop [19]
nop [19]
.wrap
This program is straightforward: it sets the pin high, waits 40 cycles, sets it low, and waits another 40 cycles, producing an LED blinking effect. Next, we need to write initialization code. The official documentation recommends embedding this directly inside the .pio file:
.program blink
.wrap_target
set pins, 1
nop [19]
nop [19]
set pins, 0
nop [19]
nop [19]
.wrap
% c-sdk {
void blink_program_init(PIO pio, uint sm, uint offset, uint pin, float div) {
pio_sm_config c = blink_program_get_default_config(offset);
pio_gpio_init(pio, pin);
pio_sm_set_consecutive_pindirs(pio, sm, pin, 1, true); // 這一行不加也可以,本範例當中只有用到 set pin
sm_config_set_set_pins(&c, pin, 1);
sm_config_set_clkdiv(&c, div);
pio_sm_init(pio, sm, offset, &c);
}
%}
It is enclosed in % c-sdk { %}:
- Fetch the default configuration via
blink_program_get_default_config(auto-generated by the build system). - Initialize the GPIO pin for PIO usage via
pio_gpio_init(pio, pin). - Set pin direction with
pio_sm_set_consecutive_pindirs(pio, sm, pin, 1, true)(falsefor input,truefor output). - Configure the SET pin mapping using
sm_config_set_set_pins(&c, pin, 1). - Configure the clock divider using
sm_config_set_clkdiv(&c, div). - Initialize the state machine with
pio_sm_init(pio, sm, offset, &c).
In addition to sm_config_set_set_pins, there are similar helpers like sm_config_set_sideset_pins.
To use PIO features, you also need to link hardware_pio in target_link_libraries:
cmake_minimum_required(VERSION 3.12)
include($ENV{PICO_SDK_PATH}/external/pico_sdk_import.cmake)
include($ENV{PICO_SDK_PATH}/tools/CMakeLists.txt)
project(pio C CXX ASM)
set(CMAKE_C_STANDARD 11)
set(CMAKE_CXX_STANDARD 17)
pico_sdk_init()
add_executable(${PROJECT_NAME}
main.c
)
pico_add_extra_outputs(${PROJECT_NAME})
pico_generate_pio_header(${PROJECT_NAME}
${CMAKE_CURRENT_LIST_DIR}/blink.pio
)
+ target_link_libraries(${PROJECT_NAME}
+ pico_stdlib
+ hardware_pio
+)
pico_enable_stdio_usb(pio 1)
pico_enable_stdio_uart(pio 1)
Next, in main.c:
#include <stdio.h>
#include "pico/stdlib.h"
#include "hardware/pio.h"
#include "hardware/clocks.h"
#include "blink.pio.h" // 編譯 pio 後自動產生 header 檔
void blink(PIO pio, uint sm, uint offset, uint pin, uint freq);
int main()
{
stdio_init_all();
PIO pio = pio0;
uint offset = pio_add_program(pio, &blink_program);
blink(pio, 0, offset, 2, 2000);
while (true)
{
printf("test");
sleep_ms(200);
}
}
void blink(PIO pio, uint sm, uint offset, uint pin, uint freq)
{
float div = clock_get_hz(clk_sys) / freq;
blink_program_init(pio, sm, offset, pin, div); // 在 PIO 裡頭宣告的函數
pio_sm_set_enabled(pio, sm, true);
}
To access clock information and PIO operations, include hardware/pio.h and hardware/clocks.h. After compiling and uploading the code, the LED blinks while “test” strings continuously print to the serial console, demonstrating that the two operations are completely decoupled and non-interfering!
Conclusion
This post provided an introductory overview of PIO and its assembly syntax. In the next post, I’ll attempt to implement a common communication protocol like UART using PIO, or interface with a DHT11 temperature sensor to deepen our understanding of PIO. To me, PIO is a truly novel concept—as far as I know, this is the first microcontroller to take this specific architectural approach.
Beyond offloading heavy protocol processing from the main CPU, this design frees developers from the limitations of fixed hardware peripherals, letting them roll their own interfaces directly via PIO. I’m very excited about the potential applications this unlocks. Furthermore, Raspberry Pi sells the RP2040 chip standalone, and you can even design your own custom board following their official documentation3.
The official documentation4 is remarkably thorough and well worth reading—it gives you a clear sense of why the hardware was designed this way. If you have questions about specific SDK functions, you can check them out here.
Footnotes
-
/devnote/2020-07-24-arduino-esp32-co2-sensor-2/ ↩
-
For development environment setup, refer to https://www.raspberrypi.com/documentation/microcontrollers/raspberry-pi-pico.html ↩
-
https://datasheets.raspberrypi.com/rp2040/hardware-design-with-rp2040.pdf ↩
-
https://datasheets.raspberrypi.com/rp2040/rp2040-datasheet.pdf ↩
Related Posts
- When a Measure Becomes a Target: From the Window Tax to Pull Request Counts I once wrote a script to tally how many PRs I contributed in a quarter, how many reviews I left, and how many tickets I closed, hoping to use numbers to prove my output to my manager. My manager simply remarked that performance isn't just about output. Years later, I finally understood—when a measure becomes a target, it ceases to be a good measure. From the British window tax and the Hanoi rat bounty to evaluating developers by PR counts today, the underlying mechanism is exactly the same.
- Using Cloudflare Images for Image Storage and Transformation Putting an image on a webpage is the simplest task in frontend development. But doing it properly—including resizing, generating multiple formats, and withstanding heavy traffic—is actually an entire end-to-end solution. Eventually, I offloaded everything to Cloudflare Images, keeping only a single original image.
- Stop Using AWS Access Keys Access Keys are an easily overlooked security risk in AWS. By pairing OIDC with IAM Roles, GitHub Actions can securely operate AWS resources without storing any secrets.
- Database Primary Keys: AUTO_INCREMENT, UUID, and UUIDv7 Backend developers often face the choice of primary keys: should you use auto-increment or UUID? What about collisions? How does UUIDv7 compare to created_at + index in performance? Here are the design decisions and benchmark results from testing 20 million rows.