Repository navigation
Releases: isarandi/poseviz
Release list
v0.5.0 — Error Reporting and EGL-First Headless
Highlights
- Failures that make a requested output impossible are no longer silently swallowed: they are raised in the caller's process as
poseviz.VisualizerErrorat the next API call, carrying the visualizer-side traceback. Previously, a broken video writer produced a clean exit 0 with no output file. - Headless mode now obtains a standalone EGL context first (invisible window only as fallback): it works regardless of
DISPLAY, targets the NVIDIA GPU on hybrid-GPU machines, and the spurious GLFW X11 warning is gone. close()returns per-sequenceposeviz.SequenceReportobjects (video path, frames written, error if any).
New features
- New exports:
poseviz.VisualizerError,poseviz.SequenceReport. - New
egl_device_indexparameter selects the EGL device for headless rendering; by default, devices are searched for an NVIDIA GPU whengpu_encode=True. - Fail-fast GPU probe: starting a video output with
gpu_encode=Trueon a non-NVIDIA GL context raises immediately with the remedies spelled out (PRIME render offload env vars, headless mode, orgpu_encode=False) instead of failing on every frame. - Sequence accounting: each output sequence is finalized and reported exactly once — whether ended explicitly, implicitly by the next
new_sequence_output(), or byclose()— and a "wrote N frames" summary is logged per successful video. - GPU-frames fast path: when downscaling keeps the frame size and the camera has no distortion, the GPU reproject collapses to a single copy.
- Four new smoke tests cover the error contract (18 total).
Behavior changes
- Writer and camera-trajectory failures poison the current sequence and raise
VisualizerErrorat the next API call; starting a new sequence resets the state, so batch loops can catch per video and continue. - A requested video output to which zero frames were written raises.
- Persistent per-frame errors escalate: 5 consecutive scene-update failures (or 10 consecutive dropped frames in the undistortion stage) while an output is active become sequence-fatal. Without an active output, per-frame errors are still only logged and visualization stays best-effort.
- An unexpected visualizer process death raises
VisualizerErroron close instead of only logging. - Context-manager exit never masks an exception from the body: pending visualizer errors are raised from
close()only on clean exit. - Headless mode with
DISPLAYset no longer lands on the display GPU via an invisible window (the cause of silent NVENC failures on hybrid-GPU machines); manually clearingDISPLAYis no longer needed.
Bug fixes
- Starting a new sequence without finalizing the previous one silently discarded the previous camera trajectory; the previous sequence is now properly finalized.
demo.py(the README's installation test) was not runnable on a fresh install: it defaulted to an SMPL variant requiring a non-public data file, and a top-levelsmplfitterimport (not a poseviz dependency) broke even the basic variant. It now runs an animated basic demo by default.
Dependency and packaging changes
- Pinned
framepump>=0.4.0for the explicit GLVideoWriterbackend=parameter (selects CUDA encoding under standalone EGL contexts).
Documentation
- Headless guide rewritten for the EGL-first context; video-output guide documents failure surfacing; architecture page documents the fault channel; README gained "Error Handling" and "Headless and GPU Notes" sections; intersphinx links to the deltacamera/framepump docs fixed (previously never resolved).
v0.4.0 — Reliability Overhaul
Highlights
- Numerous deadlocks fixed:
pause()/resume(),paused=True, window closing andclose()are now reliable, and a bad frame can no longer freeze the pipeline. world_upworks again in the GL backend, and camera trajectory recording is reinstated.- matplotlib is no longer a dependency: all 170 colormaps are shipped as vendored lookup tables.
New features
world_upis functional again: camera orbit/pan/fly and the ground plane follow the configured up axis, so Z-up worlds render correctly.- Camera trajectory recording reinstated:
camera_trajectory_pathagain produces a pickled list of(frame_index, camera)tuples (same format as pre-0.3 releases), instead of an empty list. update_multiviewaccepts more views than initially allocated: extra views fall back to queue transfer with a one-time performance warning.- All matplotlib colormap names (including
_rvariants) are available for scalar mesh coloring via vendored tables. - New headless smoke test suite:
tests/smoke_test.py(14 end-to-end scenarios).
Behavior changes
update()/update_multiview()validate inputs and raiseValueError(frame without camera, empty view list) instead of silently freezing the pipeline.- Caller-provided arrays and ViewInfo objects are never mutated anymore; tuples, integer boxes and
boxes=Noneare accepted. - No video frames are written while paused (previously, duplicates were written at wall-clock rate).
resume()is signaled out-of-band; calling it when nothing is paused cancels the next pause.- Closing the window no longer ends visualization: processing continues off-screen so video output completes and the main process is not blocked.
- Unknown colormap names now raise
ValueErrorlisting the available names.
Bug fixes
Pipeline and lifecycle
pause()followed byresume(), andPoseViz(paused=True), deadlocked the pipeline.audio_pathwas consumed as the camera-trajectory path: the audio source file was overwritten with a pickle and no audio was muxed.- Errors in undistortion workers killed the feeder thread, silently stalling everything; renderer errors likewise. Both now drop the affected frame and continue.
close()hung when the visualizer process had died or when called twice; it is now idempotent.- Ring-buffer capacity was exactly at the theoretical limit; frames held outside the queues could be overwritten while displayed.
- uint16 frames used the low byte instead of the high byte; box rescaling swapped the x/y scale factors; boxes were not reprojected through undistortion on the GPU path.
Rendering
- Meshes drew the entire reserved index buffer: garbage triangles from uninitialized memory and ghost bodies after the person count dropped.
- VRAM leaked on every window resize/fullscreen toggle and for every departing body.
- Colormap and image-plane textures wrapped around, blending opposite ends/edges; textured-mesh opacity had no effect.
- The tube shader produced NaNs for coincident joints (collapsed limbs silently vanished).
- Mouse coordinates were misaligned on HiDPI displays; fullscreen toggling tracked the wrong resolution.
- CUDA-GL texture registration broke when textures were recreated at a new size (GL recycles object names).
- Stale camera pyramids stayed clickable after the view count shrank; camera indices were not clamped, crashing on shrink.
Dependency and packaging changes
- matplotlib removed as a runtime dependency (previously imported but never declared — a fresh install was broken); colormap tables are vendored (136 KiB, regenerable via
tools/generate_colormap_luts.py). - Dropped unused dependencies:
more-itertools,imageio,rlemasklib. - Pinned
framepump>=0.2.0(GLVideoWriter,start_sequence(gpu=...)).
v0.3.3 — Display-Free Headless Rendering
What's New
Headless Rendering Without a Display Server
Headless mode no longer requires a running X11/Wayland display server. Previously, headless rendering obtained its OpenGL context from an invisible GLFW window, which still needed a display connection — on truly display-less machines (SSH sessions, containers, cluster nodes), GLFW initialization failed. Now, when GLFW cannot initialize or cannot create a window, the renderer falls back to a standalone EGL context via ModernGL, giving fully GPU-accelerated offscreen rendering with no display server at all.
Combined with framepump's CUDA NVENC encoder path (used automatically when DISPLAY is unset), video rendering now runs end-to-end on machines without any windowing system. When a display is available, behavior is unchanged.
v0.3.2 — ModernGL Rendering Backend
What's New in v0.3
This is a major release that completely replaces the rendering backend — from Mayavi to a custom OpenGL renderer built on ModernGL and GLFW. The new renderer is faster, more interactive, and supports GPU-accelerated video encoding.
New Rendering Engine
The entire rendering pipeline has been rewritten from scratch:
- ModernGL + GLFW: Replaces the Mayavi/VTK dependency with a lightweight OpenGL stack. Faster startup, lower memory usage, no Qt dependency.
- Instanced rendering: Skeleton joints (spheres) and limbs (tubes) are rendered with GPU instancing — one draw call per color group instead of one per joint.
- MSAA antialiasing: 4x multisampled framebuffer for smooth edges.
- Multiple mesh color sources: Body meshes support uniform color, per-vertex RGB, scalar-to-colormap mapping, and UV-mapped textures — each with its own shader variant.
- Raymond 3-point lighting: Camera-relative Lambertian lighting for consistent illumination from any angle.
Interactive Camera
The new free-fly terrain camera provides full interactive navigation:
- Orbit (left drag): Rotate around a pivot point in the scene.
- Look around (Shift + left drag): Rotate the viewing direction in place — the camera stays fixed while the view rotates. Useful for surveying a scene from a fixed vantage point.
- Pan (middle drag): Move the pivot parallel to the view plane.
- Zoom (right drag or scroll wheel): Change distance from pivot.
- Fly (arrow keys, Page Up/Down): Move through the scene in the camera's look direction.
- Field of view (+/−): Adjust the camera's FOV.
- Snap to camera (number keys 1–9): Jump to a displayed camera and track it. Any manual movement unsnaps.
- Camera history (mouse back/forward buttons): Navigate between previous camera positions, like a browser's back/forward.
Split-Screen Mode
Press Tab to toggle split-screen mode:
- Left pane: The original camera view (locked to the recording camera or scripted
viz_camera). - Right pane: The free-fly camera — fully interactive, independent of the left pane.
This lets you see exactly what the camera sees while simultaneously exploring the scene from a custom angle. Mouse and keyboard input only affects the right pane.
GPU Video Encoding
- Zero-copy NVENC encoding: The rendered framebuffer texture is passed directly to the hardware encoder via CUDA-OpenGL interop. No pixel readback to CPU.
- CPU fallback: Set
gpu_encode=Falsefor machines without NVENC support. - Independent render resolution: Set
render_resolution=(1920, 1080)to render at high resolution for video output while the window displays at a lower resolution. - Audio copying: Pass
audio_pathto copy the audio track from the source video into the output. - Mid-session recording: Start and stop recording with
new_sequence_output()andfinalize_sequence_output(). Record multiple segments to different files in a single session.
Auto-Headless Detection
The headless parameter now defaults to None (auto-detect). If neither DISPLAY nor WAYLAND_DISPLAY is set (e.g., SSH without X forwarding), PoseViz automatically switches to headless mode. The same code works on a desktop (opens a window) and on a remote server (renders offscreen to video) without changes.
GPU Frame Input
The new gpu_frames=True flag enables passing GPU tensors directly as frames, avoiding CPU round-trips:
- Accepts PyTorch CUDA tensors and any DLPack-compatible object (CuPy, JAX, hardware video decoder outputs, etc.).
- Image downscaling and undistortion run on the GPU via
deltacamera.pt.reproject_image(singlegrid_samplekernel). - Enables a fully GPU-resident pipeline: decode (NVDEC) → resize/undistort (CUDA) → render (OpenGL) → encode (NVENC).
Scripted Camera with Interactive Override
When a viz_camera is passed to update(), it drives the view. But the user can take over at any time by dragging the mouse — the terrain camera picks up from the current view position. On mouse release, the next update() with a viz_camera resumes the scripted camera path.
Documentation
New Sphinx documentation with Diátaxis structure:
- How-to guides: Usage tips, video output, headless rendering, GPU frames, multi-view visualization.
- Explanations: Architecture (process model, shared memory), coordinate systems, rendering pipeline (scene graph, instancing, lighting, camera controls, split-screen viewports).
Breaking Changes
- Mayavi is no longer a dependency. The rendering backend is entirely new. If you were using Mayavi-specific features or accessing internal Mayavi objects, those no longer exist.
cameravisionreplaced bydeltacamerafor camera intrinsics/extrinsics.torch_framesrenamed togpu_framesto reflect that any DLPack-compatible GPU object is accepted, not just PyTorch tensors.- Python ≥ 3.10 required (was ≥ 3.8).
- New dependencies:
moderngl,glfw,deltacamera,framepump,simplepyutils,rlemasklib,numba. Removed:mayavi,cameravision,boxlib.