Skip to content

Releases: isarandi/poseviz

v0.5.0 — Error Reporting and EGL-First Headless

Choose a tag to compare

@isarandi isarandi released this 08 Aug 22:11

Highlights

  • Failures that make a requested output impossible are no longer silently swallowed: they are raised in the caller's process as poseviz.VisualizerError at the next API call, carrying the visualizer-side traceback. Previously, a broken video writer produced a clean exit 0 with no output file.
  • Headless mode now obtains a standalone EGL context first (invisible window only as fallback): it works regardless of DISPLAY, targets the NVIDIA GPU on hybrid-GPU machines, and the spurious GLFW X11 warning is gone.
  • close() returns per-sequence poseviz.SequenceReport objects (video path, frames written, error if any).

New features

  • New exports: poseviz.VisualizerError, poseviz.SequenceReport.
  • New egl_device_index parameter selects the EGL device for headless rendering; by default, devices are searched for an NVIDIA GPU when gpu_encode=True.
  • Fail-fast GPU probe: starting a video output with gpu_encode=True on a non-NVIDIA GL context raises immediately with the remedies spelled out (PRIME render offload env vars, headless mode, or gpu_encode=False) instead of failing on every frame.
  • Sequence accounting: each output sequence is finalized and reported exactly once — whether ended explicitly, implicitly by the next new_sequence_output(), or by close() — and a "wrote N frames" summary is logged per successful video.
  • GPU-frames fast path: when downscaling keeps the frame size and the camera has no distortion, the GPU reproject collapses to a single copy.
  • Four new smoke tests cover the error contract (18 total).

Behavior changes

  • Writer and camera-trajectory failures poison the current sequence and raise VisualizerError at the next API call; starting a new sequence resets the state, so batch loops can catch per video and continue.
  • A requested video output to which zero frames were written raises.
  • Persistent per-frame errors escalate: 5 consecutive scene-update failures (or 10 consecutive dropped frames in the undistortion stage) while an output is active become sequence-fatal. Without an active output, per-frame errors are still only logged and visualization stays best-effort.
  • An unexpected visualizer process death raises VisualizerError on close instead of only logging.
  • Context-manager exit never masks an exception from the body: pending visualizer errors are raised from close() only on clean exit.
  • Headless mode with DISPLAY set no longer lands on the display GPU via an invisible window (the cause of silent NVENC failures on hybrid-GPU machines); manually clearing DISPLAY is no longer needed.

Bug fixes

  • Starting a new sequence without finalizing the previous one silently discarded the previous camera trajectory; the previous sequence is now properly finalized.
  • demo.py (the README's installation test) was not runnable on a fresh install: it defaulted to an SMPL variant requiring a non-public data file, and a top-level smplfitter import (not a poseviz dependency) broke even the basic variant. It now runs an animated basic demo by default.

Dependency and packaging changes

  • Pinned framepump>=0.4.0 for the explicit GLVideoWriter backend= parameter (selects CUDA encoding under standalone EGL contexts).

Documentation

  • Headless guide rewritten for the EGL-first context; video-output guide documents failure surfacing; architecture page documents the fault channel; README gained "Error Handling" and "Headless and GPU Notes" sections; intersphinx links to the deltacamera/framepump docs fixed (previously never resolved).

v0.4.0 — Reliability Overhaul

Choose a tag to compare

@isarandi isarandi released this 23 Jul 21:59

Highlights

  • Numerous deadlocks fixed: pause()/resume(), paused=True, window closing and close() are now reliable, and a bad frame can no longer freeze the pipeline.
  • world_up works again in the GL backend, and camera trajectory recording is reinstated.
  • matplotlib is no longer a dependency: all 170 colormaps are shipped as vendored lookup tables.

New features

  • world_up is functional again: camera orbit/pan/fly and the ground plane follow the configured up axis, so Z-up worlds render correctly.
  • Camera trajectory recording reinstated: camera_trajectory_path again produces a pickled list of (frame_index, camera) tuples (same format as pre-0.3 releases), instead of an empty list.
  • update_multiview accepts more views than initially allocated: extra views fall back to queue transfer with a one-time performance warning.
  • All matplotlib colormap names (including _r variants) are available for scalar mesh coloring via vendored tables.
  • New headless smoke test suite: tests/smoke_test.py (14 end-to-end scenarios).

Behavior changes

  • update()/update_multiview() validate inputs and raise ValueError (frame without camera, empty view list) instead of silently freezing the pipeline.
  • Caller-provided arrays and ViewInfo objects are never mutated anymore; tuples, integer boxes and boxes=None are accepted.
  • No video frames are written while paused (previously, duplicates were written at wall-clock rate).
  • resume() is signaled out-of-band; calling it when nothing is paused cancels the next pause.
  • Closing the window no longer ends visualization: processing continues off-screen so video output completes and the main process is not blocked.
  • Unknown colormap names now raise ValueError listing the available names.

Bug fixes

Pipeline and lifecycle

  • pause() followed by resume(), and PoseViz(paused=True), deadlocked the pipeline.
  • audio_path was consumed as the camera-trajectory path: the audio source file was overwritten with a pickle and no audio was muxed.
  • Errors in undistortion workers killed the feeder thread, silently stalling everything; renderer errors likewise. Both now drop the affected frame and continue.
  • close() hung when the visualizer process had died or when called twice; it is now idempotent.
  • Ring-buffer capacity was exactly at the theoretical limit; frames held outside the queues could be overwritten while displayed.
  • uint16 frames used the low byte instead of the high byte; box rescaling swapped the x/y scale factors; boxes were not reprojected through undistortion on the GPU path.

Rendering

  • Meshes drew the entire reserved index buffer: garbage triangles from uninitialized memory and ghost bodies after the person count dropped.
  • VRAM leaked on every window resize/fullscreen toggle and for every departing body.
  • Colormap and image-plane textures wrapped around, blending opposite ends/edges; textured-mesh opacity had no effect.
  • The tube shader produced NaNs for coincident joints (collapsed limbs silently vanished).
  • Mouse coordinates were misaligned on HiDPI displays; fullscreen toggling tracked the wrong resolution.
  • CUDA-GL texture registration broke when textures were recreated at a new size (GL recycles object names).
  • Stale camera pyramids stayed clickable after the view count shrank; camera indices were not clamped, crashing on shrink.

Dependency and packaging changes

  • matplotlib removed as a runtime dependency (previously imported but never declared — a fresh install was broken); colormap tables are vendored (136 KiB, regenerable via tools/generate_colormap_luts.py).
  • Dropped unused dependencies: more-itertools, imageio, rlemasklib.
  • Pinned framepump>=0.2.0 (GLVideoWriter, start_sequence(gpu=...)).

v0.3.3 — Display-Free Headless Rendering

Choose a tag to compare

@isarandi isarandi released this 16 Jul 21:14

What's New

Headless Rendering Without a Display Server

Headless mode no longer requires a running X11/Wayland display server. Previously, headless rendering obtained its OpenGL context from an invisible GLFW window, which still needed a display connection — on truly display-less machines (SSH sessions, containers, cluster nodes), GLFW initialization failed. Now, when GLFW cannot initialize or cannot create a window, the renderer falls back to a standalone EGL context via ModernGL, giving fully GPU-accelerated offscreen rendering with no display server at all.

Combined with framepump's CUDA NVENC encoder path (used automatically when DISPLAY is unset), video rendering now runs end-to-end on machines without any windowing system. When a display is available, behavior is unchanged.

v0.3.2 — ModernGL Rendering Backend

Choose a tag to compare

@isarandi isarandi released this 21 Mar 11:35

What's New in v0.3

This is a major release that completely replaces the rendering backend — from Mayavi to a custom OpenGL renderer built on ModernGL and GLFW. The new renderer is faster, more interactive, and supports GPU-accelerated video encoding.

New Rendering Engine

The entire rendering pipeline has been rewritten from scratch:

  • ModernGL + GLFW: Replaces the Mayavi/VTK dependency with a lightweight OpenGL stack. Faster startup, lower memory usage, no Qt dependency.
  • Instanced rendering: Skeleton joints (spheres) and limbs (tubes) are rendered with GPU instancing — one draw call per color group instead of one per joint.
  • MSAA antialiasing: 4x multisampled framebuffer for smooth edges.
  • Multiple mesh color sources: Body meshes support uniform color, per-vertex RGB, scalar-to-colormap mapping, and UV-mapped textures — each with its own shader variant.
  • Raymond 3-point lighting: Camera-relative Lambertian lighting for consistent illumination from any angle.

Interactive Camera

The new free-fly terrain camera provides full interactive navigation:

  • Orbit (left drag): Rotate around a pivot point in the scene.
  • Look around (Shift + left drag): Rotate the viewing direction in place — the camera stays fixed while the view rotates. Useful for surveying a scene from a fixed vantage point.
  • Pan (middle drag): Move the pivot parallel to the view plane.
  • Zoom (right drag or scroll wheel): Change distance from pivot.
  • Fly (arrow keys, Page Up/Down): Move through the scene in the camera's look direction.
  • Field of view (+/−): Adjust the camera's FOV.
  • Snap to camera (number keys 1–9): Jump to a displayed camera and track it. Any manual movement unsnaps.
  • Camera history (mouse back/forward buttons): Navigate between previous camera positions, like a browser's back/forward.

Split-Screen Mode

Press Tab to toggle split-screen mode:

  • Left pane: The original camera view (locked to the recording camera or scripted viz_camera).
  • Right pane: The free-fly camera — fully interactive, independent of the left pane.

This lets you see exactly what the camera sees while simultaneously exploring the scene from a custom angle. Mouse and keyboard input only affects the right pane.

GPU Video Encoding

  • Zero-copy NVENC encoding: The rendered framebuffer texture is passed directly to the hardware encoder via CUDA-OpenGL interop. No pixel readback to CPU.
  • CPU fallback: Set gpu_encode=False for machines without NVENC support.
  • Independent render resolution: Set render_resolution=(1920, 1080) to render at high resolution for video output while the window displays at a lower resolution.
  • Audio copying: Pass audio_path to copy the audio track from the source video into the output.
  • Mid-session recording: Start and stop recording with new_sequence_output() and finalize_sequence_output(). Record multiple segments to different files in a single session.

Auto-Headless Detection

The headless parameter now defaults to None (auto-detect). If neither DISPLAY nor WAYLAND_DISPLAY is set (e.g., SSH without X forwarding), PoseViz automatically switches to headless mode. The same code works on a desktop (opens a window) and on a remote server (renders offscreen to video) without changes.

GPU Frame Input

The new gpu_frames=True flag enables passing GPU tensors directly as frames, avoiding CPU round-trips:

  • Accepts PyTorch CUDA tensors and any DLPack-compatible object (CuPy, JAX, hardware video decoder outputs, etc.).
  • Image downscaling and undistortion run on the GPU via deltacamera.pt.reproject_image (single grid_sample kernel).
  • Enables a fully GPU-resident pipeline: decode (NVDEC) → resize/undistort (CUDA) → render (OpenGL) → encode (NVENC).

Scripted Camera with Interactive Override

When a viz_camera is passed to update(), it drives the view. But the user can take over at any time by dragging the mouse — the terrain camera picks up from the current view position. On mouse release, the next update() with a viz_camera resumes the scripted camera path.

Documentation

New Sphinx documentation with Diátaxis structure:

  • How-to guides: Usage tips, video output, headless rendering, GPU frames, multi-view visualization.
  • Explanations: Architecture (process model, shared memory), coordinate systems, rendering pipeline (scene graph, instancing, lighting, camera controls, split-screen viewports).

Breaking Changes

  • Mayavi is no longer a dependency. The rendering backend is entirely new. If you were using Mayavi-specific features or accessing internal Mayavi objects, those no longer exist.
  • cameravision replaced by deltacamera for camera intrinsics/extrinsics.
  • torch_frames renamed to gpu_frames to reflect that any DLPack-compatible GPU object is accepted, not just PyTorch tensors.
  • Python ≥ 3.10 required (was ≥ 3.8).
  • New dependencies: moderngl, glfw, deltacamera, framepump, simplepyutils, rlemasklib, numba. Removed: mayavi, cameravision, boxlib.

v0.2.1

Choose a tag to compare

@isarandi isarandi released this 21 May 22:15
Add docs, update for mesh support, PyPI