Posts

WIP: Swindle CH32V003 support

Image
 It has been a long time this has been on my todo list. I made the sw breakpoint framework a while back because i knew i would need it for that. Same for the RISCV flashstub framework. So with the assistance of deepseek/gemini to speed things up : Load and debug CH32V003 chip And it works, with breakpoints and everything : It is still early and a bit slow. Important :  - You ABSOLUTELY need a 1kB pullup on the IO pin (same pin used for SWDIO or RVIO). That should not be a problem when debugging CH32V23 or Arm chip. - Only RP2040 host for the moment

ESP32-Swindle 0.3 is out

Image
 Now that the new esp idf and associated rust binding (esp-idf, esp-svc,..) are out,  time to update esp32-swindle with S3, C3 and C6 support You'll need to use the Espressif BLE provisioning tool to transfer the wifi credentials. (and use the console to get the IP address at least once) NB: The zero boards usually have pretty bad WIFI connection An ultra cheap ESP32C3 debugging a STM32 bluepill over wifi :  Github link

RISCV memory access

 This is a small summary of why accessing memory when debugging a riscv cpu is complicated. It is work in progress, the current code is available on swindle's github. Introduction As usual, as it is RISCV, a lot is optional with important details left to implementation So let's see what is possible (definition from Gemini): Recommended : Program Buffer Execution (ProgBuf): Memory is accessed by halting the CPU and executing standard load or store instructions placed into a dedicated buffer. This uses the core's native pipeline, meaning all MMU translations, Physical Memory Protection (PMP), and cache hierarchy rules apply automatically. Optional: System Bus Access (SysBus): The Debug Module acts as an independent master directly on the system interconnect, bypassing the CPU entirely. This allows reading or writing physical memory while the core is actively running, though it does not go through the CPU's MMU or local L1/L2 caches. (NB this is more or less what ARM doe...

CH32v3xx internal usb bootloader

 That's a topic,  i've been playing with recently. The CH32V3xx and CH32V2xx RISC-V families have a internal bootloader that is able to flash software over USB or UART. No need for swindle, wch-link, ... (you'll need them to debug though). You can use that to flash the CH32V3xx DFU bootloader for example. There is a very nice project   CHISP-Flasher  that can flash elf/bin/.. over USB and available both as CLI or with a Qt6 GUI. Warning: not only do you need to have the BOOT0 pin High at boot, but you ALSO need to have the PB2  pin grounded at boot.

Swindle 0.7 out

  Faster bootloaders for GD32 & CH32 (you can keep using the old ones if you have them already): - Fix(?) flashing issue on RP2040 - Faster CH32V2/3 flashing - Better FreeRTOS support for ARM, Riscv & WCH - Experimental support for RP2040+W5500 (Ethernet swindle) - Updated Blackmagic engine -  General stability fixes Download link

Swindle : Downloading code to the CH32VXX faster

 Now that the new bootloader has shown that it can be fast to write to the CH32v3xx flash, time to revisit the swindle way of loading code. (Gemini was really helpful again. But as a smart assistant, no vibe coding here). Starting point: So we started around average speed of 3kB/S (including erase, download and write) Optim 1: use bigger blocks Before, we were configuring the block size/erase size to be 256 bytes (same as the small hardware pages). Reconfiguring the description to be 4kB gave a small speed bump => 5 kB (internally it will call 256 bytes multiple times, it just reduces the overhead). Optim 2: Faster writing to RAM The dominant slowdown is the amount of time needed to write to the CH32V3xx ram (then we call the flashstub to actually write it, that's pretty fast). Just the writing caps the overall speed to 9 kB/s , without even erasing or writing. So the optimisation is to use a fast_write_ram function there that is using abstract commands with auto exec to pipelin...

Revisiting the CH32V303 bootloader (with gemini), it is now FAST

 The DFU bootloader i wrote a while back for the swindle running on the CH32V303/CH32V307 was a quick job. It contained a cut down version of Esprit (the framework) with FreeRTOS and everything. It was about 14kB big. Since i'm playing with antigravity/gemini, it was a good sample task to run it. (On the antigravity front : it does a good job. But you really want to give small tasks, with clear instructions and  check whatever it wants to do. It tends to circumcumvent/hide problems rather than solving them.) Now back to the bootloader. The previous scheme was to use the native 256 bytes for erase and write. Write is actually fast, but erasing the whole 240kB was slow. It happens than the CH32v3xx chips have 2 extra modes compared the the bluepill family  : 32 kB and 64 kB erase. So i switched to a slightly incorrect algorithm (basically if asked to erase 4Kb at the beginning of a 32kB block address, erase the full 32 kB, and dont erase already erased blocks). And now it's...

RP2040+W5500 = Ethernet Swindle

Image
 I was not really happy with the CH32V307 Ethernet Version. It is a all-included & cheap version. But it's only 10 Mbps and there is not enough flash/sram to comfortably host swindle. The ESP32S3 wifi based version is sort of working, but the performances are not that great and i had to hack a lot to make rust + esp + cmake based project playing nice together. So here comes the new challenger : W5500 + RP2040 The RP2040 is probably the best host for swindle : - Clock accurate  SWD/RVSWD IO through PIO - Plenty of RAM/Flash - Top notch datasheet The W5500 is a nice UDP/TCP over SPI adapter. The MAC is completely managed by the chip so there is no latency bottleneck due to SPI. Since i'm in the "playing with agent" phase, i rewrote the w5500 driver with the help of deepseek and gemini to be very event driven and not polling driven. The agents  fixed a couple of subtle bugs ( & created some) It works, still need a bit of love.  Mega chain : a GD32F303 debugging ...

ESP32S3 mini

Image
 A quick port to support the ESP32S3 mini over Wifi. The pic below is a ESP32S3 debugging a bluepill board (still over wifi) Github page

Swindle Preview 6

New year version  Preview 6 https://github.com/mean00/swindle/releases Ethernet version for the CH32V307 eval board RTT & serial fixes Updated blackmagic engine Breakpoints in Flash for CH32V1/2 

Ethernet Swindle and gcc vs clang

Image
 There will be soon a new flavor of the swindle : The ethernet swindle It's starting to work nicely, utltimately the goal is to use the CH32V208 that is available for ~ 4$ on aliexpress. It's basically a CH32V307 without HS usb and without fpu,  same flash, same ram (more on that later).  LWIP is pretty big and i'm running short on both flash & ram. While looking into it, i discovered a nice option that gcc has and clang does not "-msave-restore" Since the Riscv does not have stmia/ldmia style instruction, it must push and pop all registers one at a time when entering/exiting a function. That is consuming a lot of code space for short functions. The -msave-restore creates function to save/restore , all possible variants and call them instead of manually pushing/popping registers. Let's give it a try, baseline is clang + hw FPU: Clang + HW FPU    255 508   (+0kB) Clang + no FPU       257 632  (+2KB) Clang + FPU+LTO  232...

SW breakpoint for the CH32V203 (and similar)

  The WCH doc clearly  states the QingKev4 b does not have hardware breakpoints. We saw that a bit earlier. The CH32V203 among others is a QingKeV4 b . I've added support for flash software breakpoint for the CH32V[2/3]xx to swindle. It is still a bit experimental. When such software breakpoints are used, under the hood the following happens : - Identify the page where the breakpoint is ( 256 bytes page for the CH32V[2/3]xx - Read the page - Store the 16 bits where the breakpoint is - Replace those 16 bits by a "ebreak" opcode - Erase the page - Write the modified page When removing the breakpoint the same thing happens except we put back the original opcode. It is a bit slow, but usable. The main drawback is that it will speed up flash wearing a lot. The framework is in place, if and when need it will be easy to add such functions to other chips. The main difference with other flashing operations are : - we want to use the smaller size possible. Usually we aim a bit hig...

Swindle : Preview 5 out

 After a long time, Swindle preview 5 is out, small changelog - Updated to latest blackmagic engine - Preliminary support for RP2350 (host and target) - Support for voltage translators - RTT support (with auto setup) - Better cortex  register support (including trustzone ones) - Overall better stability, it should freeze much less often - and of course continued support for CH32V30x (host and target) 

Swindle : RP2350 Coming soon (as host)

Now that the prices are coming down, making your own debug probe with a RP2350 is almost there. Don't expect big changes, the RP2040 is already powerful enough. Nb: since the blackmagic already supports the RP2350, it is already included

Swindle : Voltage translator merged

Image
 The latest bunch of change was merged. The main change is the addition of a voltage translator to allow operation with a lower voltage chip. There is a jumper to select between native (3.3v) and translated (1.2,  1.5, ..) voltage Example with a RP2040 + Carrier board with voltage translators (They are a pain to solder btw, i put them too close to each others). Only CH32V3xx and RP2040 for now !

CH32V3, debug in ram with vscode continued

Summary of previous episodes (CH32v3): 1- Reallocate shadow ram to have 128k of RAM 2- In flash, put a basic harmless loop 3- Tweak the linker script to put everything in ram 4- Load the code (to ram), change the PC  to the code in ram 5- You now upload much faster with infinite software breakpoints Vscode You can use the cortexm extension with riscv, it works fine with one caveat : you cannot "attach" You can only "launch" (this is specific to riscv) and that's the root of the problem. The init script looks like this :     {              "name" : "riscv GCC (pico)" , "type" : "gdb" , "request" : "launch" , "cwd" : "${workspaceRoot}" , "target" : "${workspaceRoot}/build/swindle_bootloader_ch32v3x_GCC_DEBUG.elf" , "gdbpath" : "${config:riscv_gdb}" , "breakAfterReset...

CH32V3xx Faster development by doing everything ram

Image
  Debugging the CH32V3xx (and V2xx) The CH32V3xx chips are rather good, they are pretty fast, cheap with tons of peripherals. Swindle supports the CH32V3xx chips (CH32V2XX as well but they are less interesting imho). Debugging with them is not so great : - They completely lack watchpoint (some revisions of WCH riscv cores do have them, but not the CH32V3x  it seems) - Only 4 breakpoints - Writing to flash is really slow compared to other similar chips ( from STM or Gigadevice) Shadow Flash ? The chip has 'shadow flash' i think.  It is using slow flash ( or RRAM/MRAM)  and shadowing them with  RAM.  At reset , it copies the flash to ram and execute from there. That way you get low cost and fast execution time. The option bytes  allocates the physical ram between shadow flash area and ram area.  But the total cannot exceed the physical amount of ram, ~ 256+64 in my case. Debugging in bigger RAM Similarly to what was done with the RP2040, we can debu...

swindle (lnBMP) v0.3

 A small release of swindle, a blackmagic derivative with rust in it : Changelog (short): - Better CH32v3xx support (host and target) - Rewrote ADiv logic so that we can ... - Use RP2040 PIO hardware to drive SWD - Update to latest blackmagic and still the M ain features : - Run on GD32F303, CH32V303, RP2040 - Support ARM devices and WCH riscv devices (CH32V2xx and CH32V3xx) - Soft breakpoint to debug code in ram (Arm only for now) - Built in FreeRTOS support through  "mon fos M0|M4|RV"

Setting up vscode+bmp/vscode to debug code in ram (RP2040)

 The main problem when debugging the code in ram is that the CPU will have started to execute whatever is in flash before you catch it, including potentially clearing the ram. As a result , the setup has to be done in 2 steps : - Have a "null" program in flash that does nothing -Tweak a little bit the vscode debug startup sequence Loop in flash The idea here is to modify the very first instruction and replace it with a endless loop. When you use lnArduino that means changing  mcus/arm_rp2040/sdk_copy/crt0.S and replacing the first instruction by       b _entry_point That way, after reset, it will harmelessly loop in the flash. Cortex-debug setup {   "version": "0.2.0", "configurations": [ { "name": "RAM- pico-load", "cwd": "${workspaceFolder}", "svdFile" : "${workspaceRoot}/.vscode/rp2040.svd", "executable": "build/st7789.elf", "gdbPath" : "${co...

CH32V203C8xx : no breakpoint ? :(

 A couple of months back i bought some CH32V203 on ebay.  Why ? They were cheap, they are pin to pin compatible with STM32F103Cxxxx, i can put them on bluepill board . They are fast riscv.  Why not. NB: by mistake i bought the ones with 64kB of flash :( But there is a BIG showstopper, there seems to be no hw breakpoints, only sw breakpoints.  If you have a lot of ram (like the RP2040) you can put the code in ram until it works fine and then put it in flash. That's not the case here. Having the debugger read/modify/write all the time to but sw breakpoint in flash is really a pain. In theory, the "flash" of the CH32 is actually flash copied to ram, so if there was a way to use the fake flash ram to put SW breakpoint in it, we would be good. For reference, the CH32V303 has only 4 breakpoints, and no watchpoint (i.e. triggers with the right capabilities) which is already a bit of a pain in the neck. I'll put those aside for the moment.