Power-up race conditions stop boot process
-
A couple of weeks ago we had a long power outage. My UPS ran out of power and the UPS monitor shut down my Netgate 4200 perfectly. Yay.
When the power returned, the Netgate, the core switch and my two WANs (T-Mobile and Starlink) all powered up at the same time.
When everything should have been back up, there was no routing going on and no WAN access. When I plugged into the console port, I saw no menu or anything. I press Enter and pfSense responded with 'booting.' It booted normally.
I first made this change: Added /boot/loader.conf.local --> autoboot_delay="1" on the hypothesis that the boot menu when booting got a stray false key-press and stopped the boot.
But, I tried a test at a later time and got the same failure to boot.
Further research suggested this is a race condition outcome due to the core switch and possibly the WAN(s) not being booted before pfSense.
So I set the following:
Set Tunable: kern.cam.boot_delay Value: 120000 2 minutes to delay start-up while the core switch, Starlink and TMHI gateways boot. Set Interfaces>TMHI and Starlink>DHCP Client Config advanced settings> Timeout: Set to 180 (wait up to 3 minutes for a response before giving up). Retry: Set to 15.A power-off of all devices test today successfully booted all the way into production with no human intervention.
I'm surprised to not find any issues with this in my search results for the forum. Maybe I didn't search well enough. But this kind of thing must happen almost every day somewhere. Or maybe I have just the right kind of equipment to be a problem. The core switch is a tp-Link SG-2008P.
As is, though, perhaps Netgate should have the kern.cam.boot_delay value configurable in the GUI (technically I used the System>Advanced>System Tunables interface, but that's not quite the same thing...) and a discussion in the manual (unless I missed it) on this kind of issue.
If you have this kind of issue, try the above. The kern.cam.boot_delay setting may vary for your situation depending on how long it takes your infrastructure surrounding your router to boot.
-
Hmm, do you have the full boot log from when you hit enter to complete the boot? It would be interesting to see exactly where that was waiting.
-
@stephenw10 unfortunately no logs from the ups-shutdown then power-on-reboot where it hung until I connected a console and pressed enter appear to have survived. If there's somewhere specific I could check that might still have evidence, let me know.
The event was on July 15, 2026 at 14:01:47 when the ups package shut the router down, and it rebooted after I pressed enter after power on at 15:05:24.
Here's the dmesg.boot from the last boot, which was at my latest test for a power-on-boot of everything at once, which as I wrote succeded, for what it's worth:
[26.03.1-RELEASE][admin@slbctg-gw-c.slbctg.home.arpa]/root: cat /var/log/dmesg.boot Copyright (c) 1992-2026 The FreeBSD Project. Copyright (c) 1979, 1980, 1983, 1986, 1988, 1989, 1991, 1992, 1993, 1994 The Regents of the University of California. All rights reserved. FreeBSD is a registered trademark of The FreeBSD Foundation. FreeBSD 16.0-CURRENT #12 plus-RELENG_26_03_1-n256546-1d1bfd578383: Wed May 20 15:20:10 UTC 2026 root@pfsense-build-release-amd64-2.eng.atx.netgate.com:/var/jenkins/workspace/pfSense-Plus-snapshots-26_03_1-main/obj/amd64/fvF1vE9r/var/jenkins/workspace/pfSense-Plus-snapshots-26_03_1-main/sources/FreeBSD-src-plus-RELENG_26_03_1/amd64.amd64/sys/pfSense amd64 FreeBSD clang version 19.1.7 (https://github.com/llvm/llvm-project.git llvmorg-19.1.7-0-gcd708029e0b2) sysctl_register_oid: can't re-use a leaf (machdep.hwpstate_pkg_ctrl)! VT(vga): resolution 640x480 CPU: Genuine Intel(R) Atom C1110 processor (2112.00-MHz K8-class CPU) Origin="GenuineIntel" Id=0x906a4 Family=0x6 Model=0x9a Stepping=4 Features=0xbfebfbff<FPU,VME,DE,PSE,TSC,MSR,PAE,MCE,CX8,APIC,SEP,MTRR,PGE,MCA,CMOV,PAT,PSE36,CLFLUSH,DTS,ACPI,MMX,FXSR,SSE,SSE2,SS,HTT,TM,PBE> Features2=0x7ffafbbf<SSE3,PCLMULQDQ,DTES64,MON,DS_CPL,VMX,EST,TM2,SSSE3,SDBG,FMA,CX16,xTPR,PDCM,PCID,SSE4.1,SSE4.2,x2APIC,MOVBE,POPCNT,TSCDLT,AESNI,XSAVE,OSXSAVE,AVX,F16C,RDRAND> AMD Features=0x2c100800<SYSCALL,NX,Page1GB,RDTSCP,LM> AMD Features2=0x121<LAHF,ABM,Prefetch> Structured Extended Features=0x239ca7eb<FSGSBASE,TSCADJ,BMI1,AVX2,FDPEXC,SMEP,BMI2,ERMS,INVPCID,NFPUSG,PQE,RDSEED,ADX,SMAP,CLFLUSHOPT,CLWB,PROCTRACE,SHA> Structured Extended Features2=0x984007bc<UMIP,PKU,OSPKE,WAITPKG,GFNI,VAES,VPCLMULQDQ,RDPID,MOVDIRI,MOVDIR64B> Structured Extended Features3=0xfc184410<FSRM,MD_CLEAR,IBT,IBPB,STIBP,L1DFL,ARCH_CAP,CORE_CAP,SSBD> XSAVE Features=0xf<XSAVEOPT,XSAVEC,XINUSE,XSAVES> IA32_ARCH_CAPS=0x188fd6b<RDCL_NO,IBRS_ALL,SKIP_L1DFL_VME,MDS_NO,TAA_NO> VT-x: PAT,HLT,MTF,PAUSE,EPT,UG,VPID,VID,PostIntr TSC: P-state invariant, performance statistics real memory = 4294967296 (4096 MB) avail memory = 3914854400 (3733 MB) Event timer "LAPIC" quality 600 ACPI APIC Table: <ALASKA A M I > WARNING: L3 data cache covers more APIC IDs than a package (7 > 3) FreeBSD/SMP: Multiprocessor System Detected: 4 CPUs FreeBSD/SMP: 1 package(s) x 4 core(s) random: registering fast source Intel Secure Key Seed random: fast provider: "Intel Secure Key Seed" random: unblocking device. ioapic0 <Version 2.0> irqs 0-119 Launching APs: 1 3 2 TCP_ratelimit: Is now initialized ipw_bss: You need to read the LICENSE file in /usr/share/doc/legal/intel_ipw.LICENSE. ipw_bss: If you agree with the license, set legal.intel_ipw.license_ack=1 in /boot/loader.conf. module_register_init: MOD_LOAD (ipw_bss_fw, 0xffffffff8078b120, 0) error 1 ipw_ibss: You need to read the LICENSE file in /usr/share/doc/legal/intel_ipw.LICENSE. ipw_ibss: If you agree with the license, set legal.intel_ipw.license_ack=1 in /boot/loader.conf. module_register_init: MOD_LOAD (ipw_ibss_fw, 0xffffffff8078b1d0, 0) error 1 ipw_monitor: You need to read the LICENSE file in /usr/share/doc/legal/intel_ipw.LICENSE. ipw_monitor: If you agree with the license, set legal.intel_ipw.license_ack=1 in /boot/loader.conf. module_register_init: MOD_LOAD (ipw_monitor_fw, 0xffffffff8078b280, 0) error 1 iwi_bss: You need to read the LICENSE file in /usr/share/doc/legal/intel_iwi.LICENSE. iwi_bss: If you agree with the license, set legal.intel_iwi.license_ack=1 in /boot/loader.conf. module_register_init: MOD_LOAD (iwi_bss_fw, 0xffffffff807aa980, 0) error 1 iwi_ibss: You need to read the LICENSE file in /usr/share/doc/legal/intel_iwi.LICENSE. iwi_ibss: If you agree with the license, set legal.intel_iwi.license_ack=1 in /boot/loader.conf. module_register_init: MOD_LOAD (iwi_ibss_fw, 0xffffffff807aaa30, 0) error 1 iwi_monitor: You need to read the LICENSE file in /usr/share/doc/legal/intel_iwi.LICENSE. iwi_monitor: If you agree with the license, set legal.intel_iwi.license_ack=1 in /boot/loader.conf. module_register_init: MOD_LOAD (iwi_monitor_fw, 0xffffffff807aaae0, 0) error 1 random: entropy device external interface wlan: mac acl policy registered kbd0 at kbdmux0 WARNING: Device "spkr" is Giant locked and may be deleted before FreeBSD 16.0. efirtc0: <EFI Realtime Clock> efirtc0: registered as a time-of-day clock, resolution 1.000000s netgate0: <Netgate 4200> netgate0: version: 0.2 vtvga0: <VT VGA driver> smbios0: <System Management BIOS> at iomem 0x61c66000-0x61c66017 smbios0: Entry point: v3 (64-bit), Version: 3.5 acpi0: <ALASKA A M I > acpi_ec1: <Embedded Controller: GPE 0x6e, ECDT> port 0x62,0x66 on acpi0 acpi0: Power Button (fixed) hpet0: <High Precision Event Timer> iomem 0xfed00000-0xfed003ff on acpi0 Timecounter "HPET" frequency 19200000 Hz quality 950 Event timer "HPET" frequency 19200000 Hz quality 550 Event timer "HPET1" frequency 19200000 Hz quality 440 Event timer "HPET2" frequency 19200000 Hz quality 440 Event timer "HPET3" frequency 19200000 Hz quality 440 Event timer "HPET4" frequency 19200000 Hz quality 440 atrtc1: <AT realtime clock> on acpi0 atrtc1: Warning: Couldn't map I/O. atrtc1: registered as a time-of-day clock, resolution 1.000000s Event timer "RTC" frequency 32768 Hz quality 0 attimer0: <AT timer> port 0x40-0x43,0x50-0x53 irq 0 on acpi0 Timecounter "i8254" frequency 1193182 Hz quality 0 Event timer "i8254" frequency 1193182 Hz quality 100 Timecounter "ACPI-fast" frequency 3579545 Hz quality 900 acpi_timer0: <24-bit timer at 3.579545MHz> port 0x1808-0x180b on acpi0 pcib0: <ACPI Host-PCI bridge> port 0xcf8-0xcff on acpi0 pci0: <ACPI PCI bus> on pcib0 pcib1: <ACPI PCI-PCI bridge> at device 6.0 on pci0 pci1: <ACPI PCI bus> on pcib1 nvme0: <Generic NVMe Device> mem 0x67c00000-0x67c03fff at device 0.0 on pci1 xhci0: <Intel Alder Lake-P Thunderbolt 4 USB controller> mem 0x4000040000-0x400004ffff at device 13.0 on pci0 xhci0: 32 bytes context size, 64-bit DMA xhci0: xECP capabilities <PROTO,PROTO,VEND(c0),LEGACY,VEND(c6),VEND(c7),VEND(c2),DEBUG,VEND(c3),VEND(d1),VEND(ce),VEND(c8),VEND(c9),VEND(ca),VEND(cc),VEND(cd),VEND(d2),VEND(cf),VEND(d3)> usbus0 on xhci0 usbus0: 5.0Gbps Super Speed USB v3.0 pci0: <simple comms, UART> at device 18.0 (no driver attached) pci0: <mass storage> at device 18.7 (no driver attached) xhci1: <Intel Alder Lake USB 3.2 controller> mem 0x4000020000-0x400002ffff at device 20.0 on pci0 xhci1: 32 bytes context size, 64-bit DMA xhci1: xECP capabilities <PROTO,PROTO,VEND(c0),LEGACY,VEND(c6),VEND(c7),VEND(c2),DEBUG,VEND(c3),VEND(c4),VEND(ce),VEND(c8),VEND(c9),VEND(ca),VEND(cb),VEND(cc),VEND(cd)> usbus1 on xhci1 usbus1: 5.0Gbps Super Speed USB v3.0 pci0: <memory, RAM> at device 20.2 (no driver attached) pci0: <serial bus> at device 21.0 (no driver attached) pci0: <serial bus> at device 21.1 (no driver attached) pci0: <simple comms> at device 22.0 (no driver attached) pci0: <serial bus> at device 25.0 (no driver attached) pci0: <serial bus> at device 25.1 (no driver attached) pcib2: <ACPI PCI-PCI bridge> at device 28.0 on pci0 pci2: <ACPI PCI bus> on pcib2 pcib3: <ACPI PCI-PCI bridge> at device 28.4 on pci0 pci3: <ACPI PCI bus> on pcib3 igc0: <Intel(R) Ethernet Controller I226-V> mem 0x67a00000-0x67afffff,0x67b00000-0x67b03fff at device 0.0 on pci3 igc0: EEPROM V2.17-0 eTrack 0x80000303 igc0: Using 1024 TX descriptors and 1024 RX descriptors igc0: Using 1 RX queues 1 TX queues igc0: Using MSI-X interrupts with 2 vectors igc0: Ethernet address: 90:ec:77:91:2d:86 igc0: netmap queues/slots: TX 1/1024, RX 1/1024 pcib4: <ACPI PCI-PCI bridge> at device 28.5 on pci0 pci4: <ACPI PCI bus> on pcib4 igc1: <Intel(R) Ethernet Controller I226-V> mem 0x67700000-0x677fffff,0x67800000-0x67803fff at device 0.0 on pci4 igc1: EEPROM V2.17-0 eTrack 0x80000303 igc1: Using 1024 TX descriptors and 1024 RX descriptors igc1: Using 1 RX queues 1 TX queues igc1: Using MSI-X interrupts with 2 vectors igc1: Ethernet address: 90:ec:77:91:2d:85 igc1: netmap queues/slots: TX 1/1024, RX 1/1024 pcib5: <ACPI PCI-PCI bridge> at device 28.6 on pci0 pci5: <ACPI PCI bus> on pcib5 igc2: <Intel(R) Ethernet Controller I226-V> mem 0x67400000-0x674fffff,0x67500000-0x67503fff at device 0.0 on pci5 igc2: EEPROM V2.17-0 eTrack 0x80000303 igc2: Using 1024 TX descriptors and 1024 RX descriptors igc2: Using 1 RX queues 1 TX queues igc2: Using MSI-X interrupts with 2 vectors igc2: Ethernet address: 90:ec:77:91:2d:84 igc2: netmap queues/slots: TX 1/1024, RX 1/1024 pcib6: <ACPI PCI-PCI bridge> at device 28.7 on pci0 pci6: <ACPI PCI bus> on pcib6 igc3: <Intel(R) Ethernet Controller I226-V> mem 0x67100000-0x671fffff,0x67200000-0x67203fff at device 0.0 on pci6 igc3: EEPROM V2.17-0 eTrack 0x80000303 igc3: Using 1024 TX descriptors and 1024 RX descriptors igc3: Using 1 RX queues 1 TX queues igc3: Using MSI-X interrupts with 2 vectors igc3: Ethernet address: 90:ec:77:91:2d:83 igc3: netmap queues/slots: TX 1/1024, RX 1/1024 isab0: <PCI-ISA bridge> at device 31.0 on pci0 isa0: <ISA bus> on isab0 pci0: <serial bus> at device 31.5 (no driver attached) acpi_acad0: <AC Adapter> on acpi0 uart2: <16750 or compatible> iomem 0xfe03e000-0xfe03e007 irq 16 on acpi0 uart2: console (115200,n,8,1) acpi_button0: <Sleep Button> on acpi0 cpu0: <ACPI CPU> on acpi0 acpi_wmi0: <ACPI-WMI mapping> on acpi0 acpi_wmi0: cannot find EC device acpi_wmi0: Embedded MOF found acpi_wmi1: <ACPI-WMI mapping> on acpi0 acpi_wmi1: cannot find EC device acpi_wmi1: Embedded MOF found acpi_spmc0: <Low Power S0 Idle (DSM sets 0x1)> on acpi0 acpi_tz0: <Thermal Zone> on acpi0 acpi_netgate_mec0: <Netgate MECx LED/GPIO controller SILC1001> on acpi0 atrtc0: <AT realtime clock> at port 0x70 irq 8 on isa0 atrtc0: Warning: Couldn't map I/O. atrtc0: registered as a time-of-day clock, resolution 1.000000s atrtc0: Can't map interrupt. est0: <Enhanced SpeedStep Frequency Control> on cpu0 cpufreq0: <CPU frequency control> on cpu0 cpufreq1: <CPU frequency control> on cpu1 cpufreq2: <CPU frequency control> on cpu2 cpufreq3: <CPU frequency control> on cpu3 Timecounter "TSC" frequency 2112003680 Hz quality 1000 Timecounters tick every 1.000 msec ugen0.1: <Intel XHCI root HUB> at usbus0 ugen1.1: <Intel XHCI root HUB> at usbus1 uhub0 on usbus0 uhub0: <Intel XHCI root HUB, class 9/0, rev 3.00/1.00, addr 1> on usbus0 ZFS filesystem version: 5 ZFS storage pool version: features support (5000) uhub1 on usbus1 uhub1: <Intel XHCI root HUB, class 9/0, rev 3.00/1.00, addr 1> on usbus1 nvme0: Allocated 64MB host memory buffer nvme_sim0: <nvme cam> on nvme0 nda0 at nvme0 bus 0 scbus0 target 0 lun 1 nda0: <Samsung SSD 980 250GB 3B4QFXO7 S64BNU0WB05003M> nda0: Serial Number S64BNU0WB05003M nda0: nvme version 1.4 nda0: 238475MB (488397168 512 byte sectors) Trying to mount root from zfs:pfSense/ROOT/Default []... uhub0: 3 ports with 3 removable, self powered uhub1: 16 ports with 16 removable, self powered ugen1.2: <American Power Conversion Back-UPS XS 1000M FW:945.d10 .D USB FW:d10> at usbus1 Root mount waiting for: usbus1 ugen1.3: <Generic Ultra Fast Media> at usbus1 umass0 on uhub1 umass0: <Generic Ultra Fast Media, class 0/0, rev 2.00/1.98, addr 2> on usbus1 da0 at umass-sim0 bus 0 scbus1 target 0 lun 0 da0: <Generic Ultra HS-COMBO 1.98> Removable Direct Access SCSI device da0: Serial Number 000000225001 da0: 40.000MB/s transfers da0: 14952MB (30621696 512 byte sectors) da0: quirks=0x2<NO_6_BYTE> (da0:umass-sim0:0:0:0): MODE SENSE for CACHE page command failed. (da0:umass-sim0:0:0:0): Mode page 8 missing, disabling SYNCHRONIZE CACHE CPU: Genuine Intel(R) Atom C1110 processor (2112.00-MHz K8-class CPU) Origin="GenuineIntel" Id=0x906a4 Family=0x6 Model=0x9a Stepping=4 Features=0xbfebfbff<FPU,VME,DE,PSE,TSC,MSR,PAE,MCE,CX8,APIC,SEP,MTRR,PGE,MCA,CMOV,PAT,PSE36,CLFLUSH,DTS,ACPI,MMX,FXSR,SSE,SSE2,SS,HTT,TM,PBE> Features2=0x7ffafbbf<SSE3,PCLMULQDQ,DTES64,MON,DS_CPL,VMX,EST,TM2,SSSE3,SDBG,FMA,CX16,xTPR,PDCM,PCID,SSE4.1,SSE4.2,x2APIC,MOVBE,POPCNT,TSCDLT,AESNI,XSAVE,OSXSAVE,AVX,F16C,RDRAND> AMD Features=0x2c100800<SYSCALL,NX,Page1GB,RDTSCP,LM> AMD Features2=0x121<LAHF,ABM,Prefetch> Structured Extended Features=0x239ca7eb<FSGSBASE,TSCADJ,BMI1,AVX2,FDPEXC,SMEP,BMI2,ERMS,INVPCID,NFPUSG,PQE,RDSEED,ADX,SMAP,CLFLUSHOPT,CLWB,PROCTRACE,SHA> Structured Extended Features2=0x984007bc<UMIP,PKU,OSPKE,WAITPKG,GFNI,VAES,VPCLMULQDQ,RDPID,MOVDIRI,MOVDIR64B> Structured Extended Features3=0xfc184410<FSRM,MD_CLEAR,IBT,IBPB,STIBP,L1DFL,ARCH_CAP,CORE_CAP,SSBD> XSAVE Features=0xf<XSAVEOPT,XSAVEC,XINUSE,XSAVES> IA32_ARCH_CAPS=0x1580fd6b<RDCL_NO,IBRS_ALL,SKIP_L1DFL_VME,MDS_NO,TAA_NO> VT-x: PAT,HLT,MTF,PAUSE,EPT,UG,VPID,VID,PostIntr TSC: P-state invariant, performance statistics [26.03.1-RELEASE][admin@slbctg-gw-c.slbctg.home.arpa]/root:and bios:
[26.03.1-RELEASE][admin@slbctg-gw-c.slbctg.home.arpa]/root: dmidecode -t bios # dmidecode 3.7 # SMBIOS entry point at 0x61c66000 Found SMBIOS entry point in EFI, reading table from /dev/mem. SMBIOS 3.5.0 present. Handle 0x0000, DMI type 0, 26 bytes Platform Firmware Information Vendor: American Megatrends International, LLC. Version: 8.01 Release Date: 02/01/2024 Address: 0xF0000 Runtime Size: 64 KiB ROM Size: 0 MiB Characteristics: PCI is supported Firmware is upgradeable Firmware shadowing is allowed Boot from CD is supported Selectable boot is supported Firmware ROM is socketed EDD is supported Japanese floppy for NEC 9800 1.2 MB is supported (int 13h) Japanese floppy for Toshiba 1.2 MB is supported (int 13h) 5.25"/360 kB floppy services are supported (int 13h) 5.25"/1.2 MB floppy services are supported (int 13h) 3.5"/720 kB floppy services are supported (int 13h) 3.5"/2.88 MB floppy services are supported (int 13h) Print screen service is supported (int 5h) 8042 keyboard services are supported (int 9h) Serial services are supported (int 14h) Printer services are supported (int 17h) CGA/mono video services are supported (int 10h) ACPI is supported USB legacy is supported BIOS boot specification is supported Targeted content distribution is supported UEFI is supported Platform Firmware Revision: 5.26 Embedded Controller Firmware Revision: 3.2 Handle 0x000C, DMI type 13, 22 bytes Firmware Language Information Language Description Format: Long Installable Languages: 1 en|US|iso8859-1 Currently Installed Language: en|US|iso8859-1 Handle 0x004F, DMI type 13, 22 bytes Firmware Language Information Language Description Format: Abbreviated Installable Languages: 1 enUS Currently Installed Language: enUS [26.03.1-RELEASE][admin@slbctg-gw-c.slbctg.home.arpa]/root: -
Hmm, nothing unexpected there. Can you replicate the issue?
-
@Mission-Ghost said in Power-up race conditions stop boot process:
When the power returned
And the power comes back for pfSense and the device connected to the console cable at the same moment ?
Then : maybe : If on of these two devices generated a stray 'send character' during init/power up, pfSense intercepts this char, and the OS finds in in his serial input and will use it as a "this is the character you've hit to abort the auto boot', so it abort and await further instruction ... silently waiting at the >boot prompt for you to communicate with it.Imho, the best way to circumvent this rather rare situation : don't leave the console connected during boot ? if you are present, remember to open the serial connection and say 'go ahead' if needed. If you are not there, the 'no console cable' will deal with the issue ^^
I've myself the console cable always hooked up to a PC, also connected on the same UPS as pfSense. Never had any issues, and I'm probably lucky. That said, the power going down for the UPS to shut down the devices, that happens not even ones a year.
-
@Gertjan the console cable is disconnected except when I use the console. So if there was a stray key press received, and I have no clear evidence to support that, the router received it into the empty console plug socket directly or internally.
During my first experiment to figure this out after the initial problem, I reproduced the boot hang simply by powering on the router, attached core switch and the T-Mobile gateway at once just as if a power failure had ended. (They are all on the same power strip from the same UPS.)
This appears to support my hypothesis that the stray keypress was not the problem, but a race condition.
-
@stephenw10 yes, I was able to replicate the issue in an experiment after the problem first arose following a long power failure that drained the UPS to a shutdown.
The router, T-Mobile gateway and core switch are all on the same UPS and power strip. For the experiment I shut down the router as the UPS package would have done, powered everything off at the power strip, waited about 10 seconds, then powered it back on at the power strip. The router hung up again on boot.
After making the adjustments to parameters discussed in my initial post, repeating the experiment resulted in the router booting into production without intervention.
-
@Mission-Ghost said in Power-up race conditions stop boot process:
The router hung up again on boot.
Could you see exactly where it stopped? Or can you repeat that and log the console output to get that info?
The fact it showed 'booting' when you pressed return is odd. That does appear in a normal boot at the console. It implies something abnormal is happening. Knowing where that appears in the boot would narrow down the possible causes significantly.
Is your 4200 running SlimBootLoader or AMI BIOS?
-
@stephenw10 it only responded 'booting' during the first event which followed the actual power failure and exhaustion of the UPS, UPS package initiated shutdown, and subsequent cold boot.
On the first experiment which successfully reproduced the issue, I did not get a 'booting' message. Upon connection of the console, I got the console menu. From there I told it to continue to boot and it did.
It would be cumbersome to reproduce this experiment now. I'd have to back out my changes or boot a previous boot environment, shut it down, repeat the experiment, the restore the b.e. with the changes, then reboot again. During a designated service time when production isn't needed. I don't get a lot of opportunities for this, but if I do, I will.
My 4200 is running the default bios that is enabled when it is delivered from the factory: I believe that's AMI. I haven't seen it in a couple of years.
-
The 4200 was switched to SBL relatively recently so it's probably AMI if it's a few years old. That's what I'm testing on here and failing to make it happen.
Curious where 'booting' might have come from.

@Mission-Ghost said in Power-up race conditions stop boot process:
Upon connection of the console, I got the console menu. From there I told it to continue to boot and it did.
Ah, do you mean the boot loader menu?
โโโ Welcome to Netgate pfSense Plus โโโโ __________________________ โ โ / ___\ โ 1. Boot Multi user [Enter] โ | /` โ 2. Boot Single user โ | / :-| โ 3. Escape to loader prompt โ | _________ ___/ /_ | โ 4. Reboot โ | /` ____ / /__ ___/ | โ 5. Cons: Serial โ | / / / / / / | โ โ | / /___/ / / / | โ Kernel: โ | / ______/ / / _ | โ 6. kernel (1 of 1) โ |/ / / / _| |_ | โ โ / /___/ |_ _| | โ Options: โ / |_| | โ 7. Boot Options โ /_________________________/ โ 8. Boot Environments โ / โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Autoboot in 0 seconds. [Space] to pauseWhat's normally refered to as the console menu is only shown when boot is already complete.
-
@stephenw10 Yes, I mean the boot loader menu. I used option 1 and it booted normally.
-
Huh so it had stopped at the boot loader menu. Usually that can only happen it the user hits a key at the console but you say it wasn't even connected?
Is it possible the 4200 was rebooting and you just happened to connect the console at that point creating some data on the input?
-
@stephenw10 it's possible, but unlikely. We'd waited long enough to expect internet service to return. When it had not, I had to collect my laptop and the console cable and go to where the router is located, plug in, set up the laptop and log in, start the terminal app and only then did I see the 'booting' message.
Recall, too, that I reproduced the frozen boot with an experiment, only this time I was presented the boot loader menu upon connection of the laptop to the router.
Also note that I no longer could reproduce the frozen boot once I had made the changes to the system parameters noted in my initial post. So the start-up delay encompassed by the parameter changes seem to have addressed the problem. Working back from that now two-minute delay to the default might yield some ideas. I don't know what the system is doing that early in the process.
Note this is a after-market self-upgraded 'plus' 4200 with a Samsung 2GB NVMe reported in good condition.
-
The kern.cam.boot_delay value waits for cam devices before trying to mount root. That's a long way past the loader menu. The dhcp client change wouldn't affect anything until after root is mounted.
It might have stopped at the mountroot> prompt if the disk wasn't ready. But that would not have continued booting just by pressing enter. -
@stephenw10 I canโt rule out a combination of factors. I also canโt rule out my recollection of the console outputs is accurate. I was focused on getting the network running again, not forensics. Itโs probably not helpful that the kernel boot log appears to be overwritten on every boot. Itโs not much of a log doing that.
Even so, Iโve experimentally shown the problem, the router failing to boot into production when powered on simultaneously with my tp link core switch (and one of two wan gateways: TMHI), has been alleviated with these parameter changes. The console outputs I saw may be red herrings or only peripherally related. Or not.
A Gemini query about why this .cam. delay might have help may be helpful in pointing the way to the root cause of pfsense not handling a boot scenario well, with the usual caveats that it may contain material misstatements of fact and erroneous conclusions:
Begin Gemini response:
Increasing
kern.cam.boot_delayaddresses this issue because it resolves a power-on race condition between your pfSense router, your core switch, and your WAN gateways during a simultaneous cold boot.
The Underlying Problem: A Boot Order Race Condition
When power is restored simultaneously across all hardware (such as after a power outage), pfSense boots up significantly faster than your surrounding network infrastructure:
- pfSense Boots Rapidly: An x86 firewall typically reaches kernel initialization and system service startup in 15 to 30 seconds.
- TP-Link Switch Initialization: Managed or smart TP-Link switches require time to complete hardware POST and transition switchports out of Spanning Tree Protocol (STP) listening/learning states (which can take 15โ30 seconds before passing traffic).
- T-Mobile Home Internet (TMHI) Boot Latency: Cellular gateways (such as Arcadyan, Sagemcom, or FAST5688W) take 1.5 to 3 minutes to power on internal Linux firmware, negotiate cellular towers, bring up Ethernet PHY link state, and start internal DHCP services.
Why pfSense Hangs
When pfSense launches its network subsystem while attached devices are still booting, several cascading failures occur:
- DHCP Client Blocking (
dhclientStalls): During boot, pfSense startsdhclienton its WAN interfaces to request an IP address. If the T-Mobile gateway's LAN ports are down or its DHCP server hasn't started,dhclientrepeatedly retries, times out, or blocks the boot script execution sequence. - Ethernet PHY Flapping: As the TP-Link switch and T-Mobile gateway power cycle their Ethernet ports during their own boot processes, physical link state toggles up and down. FreeBSD kernel network drivers (
igb,em,re) responding to rapid link-state changes during early startup can stall system scripts (like interface assignment or firewall rule compilation). - Service Dependencies: Key startup services (gateway monitoring via
dpinger, DNS resolution viaunbound, NTP, and NAT configuration) depend on WAN interface initialization. If interface queries fail or hang, system startup scripts pause or freeze indefinitely waiting for network responses that aren't coming.
How
kern.cam.boot_delaySolves Itkern.cam.boot_delayis a FreeBSD kernel tunable (measured in milliseconds). Although CAM (Common Access Method) is FreeBSD's disk and storage subsystem, setting this parameter forces the kernel to pause early in the boot sequence before mounting the root filesystem and initializing network interfaces.Setting
kern.cam.boot_delay(for example, to120000for 2 minutes) introduces a intentional pause:- Gives Peripherals Time to Prepare: The pfSense kernel sits idle before bringing up network interfaces, giving the TP-Link switch time to establish link and the T-Mobile gateway time to establish cell connections and spin up its local DHCP server.
- Ensures Clean Link Negotiation: When pfSense finally resumes booting and initializes its NICs, all attached Ethernet links are stable and active.
- Smooth DHCP Retrieval: The WAN interface receives an IP address on its first DHCP request without timing out or blocking execution.
Recommended Settings
To ensure reliable unattended reboots after a power outage, pair the CAM boot delay with DHCP client timeout adjustments in pfSense:
- Set Boot Delay (System Tunables or
/boot/loader.conf.local):
- Tunable Name:
kern.cam.boot_delay - Value:
60000to120000(60 to 120 seconds, depending on how long your TMHI gateway takes to hand out an IP).
- Adjust WAN DHCP Client Settings (pfSense WebGUI):
- Go to Interfaces > [TMHI WAN].
- Under DHCP Client Configuration (Advanced Options), set Timeout to
180seconds and Retry to15seconds. This prevents pfSense from giving up on the WAN if the gateway takes slightly longer to hand out an IP address.
End Gemini response.
-
There used to be a bug in dhclient that would cause it to stop trying to pull a lease if the interfaces flapped at the wrong time but that should no longer be the case. But even when that did happen it didn't prevent the boot process completing.
Yes, adding a delay anywhere will prevent the unexpected interaction with the other devices at boot if ports link cycle.
It would still like to see a console log showing it fail to boot. If there is some bug there I'd love to squash it.
Privacy Policy · Cookie Policy