Midpilot Connector Generator for discovery, scraping, digester and Codegen built with FastAPI.
The source tree is layered by purpose and import direction: an item may import
anything below it in the list, never above. The layering is enforced by
import-linter (uv run poe importcheck, configured in pyproject.toml).
docs- documentationsrc/app.py,src/router.py- composition root (FastAPI app, route aggregation)src/modules- feature pipelines (discovery, scrape, digester, codegen)src/session- session context: routes, documentation upload/processing, ownership checksrc/auth- API key authentication and key management endpointssrc/api- shared HTTP edge: exception handlers, job status response builderssrc/jobs- durable PostgreSQL job queue (claiming worker, runner, lifecycle, caching, fencing, persistence)src/documents- documentation toolkit: chunking, LLM processing, filtering, relevancesrc/integrations- adapters for external services (web fetch and search)src/database- SQLAlchemy models and repositories (single shared schema)src/core- technical foundations: LLM client, observability, DB engine, base errors/schemasrc/shared- pure helpers and product-wide vocabulary (no imports from other src packages)src/config- settings loaded from environmenttest/unit,test/integration- tests, mirroring the src layout
Important files:
server.py- Hypercorn server entry pointsrc/app.py- FastAPI entry pointsrc/config- project configurationpyproject.toml- dependencies, tools, tasks, import-linter contractsDockerfile- docker fileDockerfile.base- reusable Python + Playwright base image
App exposes API documentation on the URLs below:
- OpenAPI UI: http://localhost:8090/docs
- ReDoc: http://localhost:8090/redoc
App is configured using environment variables, see the default configuration in src/config.py and samples in .env-example.
For local development you can take advantage of dotenv plugins that read .env file from the project root.
# copy and use example dev configuration
cp .env-example .env
# copy and use configuration for unit/integration tests
cp .env.test-example .env.testGravitee issues and validates API keys. The backend has no key-management endpoints, local key registry, or master key.
AUTH__MODE=prod(default): require oneX-Gravitee-Api-Keyheader validated and forwarded by Gravitee. Restrict backend access to the gateway; the backend does not verify issuance, expiration or revocation itself.AUTH__MODE=dev: explicitly bypass authentication and ownership checks for local development. Supplied keys are ignored and new sessions are ownerless. Never use this mode for a shared or public deployment.
Sessions and reusable job results are isolated by the SHA-256 fingerprint of that exact API key. The raw key is not persisted. A new/rotated key cannot access an old key's sessions. Ownerless development sessions are inaccessible through Gravitee mode. Supported session and pipeline response bodies are unchanged.
See Gravitee setup and local verification for gateway requirements, consumer commands, and deployment checks. Upgrading from local API keys requires removing the obsolete key table and session-owner column; follow the milestone upgrade instructions.
# build base image when Python, uv.lock, pyproject.toml, or Playwright changes
docker build -f Dockerfile.base -t midpilot-connector-gen-base:python3.13-playwright1.61.0 .
# build image
docker compose build
# start container
docker compose upNOTE: tasks are run with a poethepoet tool and configured in pyproject.toml
The application requires PostgreSQL database. The database is automatically configured when using Docker Compose.
The docker-compose.yaml includes a PostgreSQL service that is automatically configured with the environment variables from your .env file:
# Start all services (app + database)
docker compose up
# Start only the database
docker compose up dbThe database service uses the following environment variables from .env:
DATABASE__HOST- Database host (default: localhost)DATABASE__INT_PORT- Internal database port used inside Docker/networked deployments (default: 5432)DATABASE__EXT_PORT- External database port exposed to the host (default: 5433)DATABASE__USER- Database userDATABASE__PASSWORD- Database passwordDATABASE__NAME- Database name
If you prefer to run PostgreSQL outside Docker Compose:
docker run --name postgres \
-e POSTGRES_USER=user \
-e POSTGRES_PASSWORD=password \
-e POSTGRES_DB=db \
-p 5432:5432 \
-d postgres:15-alpineEnsure your .env file has the correct database configuration:
DATABASE__HOST=localhost
DATABASE__INT_PORT=5432
DATABASE__EXT_PORT=5433
DATABASE__USER=user
DATABASE__PASSWORD=password
DATABASE__NAME=db
# Full connection string for Alembic and SQLAlchemy
DATABASE__URL=postgresql+asyncpg://${DATABASE__USER}:${DATABASE__PASSWORD}@${DATABASE__HOST}:${DATABASE__EXT_PORT}/${DATABASE__NAME}The application uses Alembic for database migrations.
# Run all pending migrations
uv run alembic upgrade head
# Check current migration version
uv run alembic current
# View migration history
uv run alembic history# Create a new migration (auto-generate from models)
uv run alembic revision --autogenerate -m "description of changes"
# Upgrade to a specific revision
uv run alembic upgrade <revision_id>
# Downgrade one revision
uv run alembic downgrade -1
# Downgrade to base (drop all tables)
uv run alembic downgrade base
# Show current revision
uv run alembic current
# Show SQL without executing
uv run alembic upgrade head --sql# after pulling a version with schema changes
uv run alembic upgrade head
# APP__WORKERS can be raised when live reload is disabled; all processes
# coordinate through the same PostgreSQL job queue
uv run poe start
# access the service at http://localhost:8090
# e.g. `curl http://0.0.0.0:8090/health`poe start runs both roles in one process. To reproduce the split deployment
locally — one API that only serves requests, plus several queue consumers —
start them as separate processes. The consumers use
python -m src.jobs.runtime, which starts no HTTP server:
# 1x API, consuming nothing
JOBS__ENABLED=false APP__WORKERS=1 APP__LIVE_RELOAD=false \
uv run python server.py > /tmp/api.log 2>&1 &
# 10x worker; lower pools because each process opens its own
for i in $(seq 1 10); do
DATABASE__POOL_SIZE=3 DATABASE__MAX_OVERFLOW=3 \
uv run python -m src.jobs.runtime > /tmp/worker-$i.log 2>&1 &
doneVerify the split took effect — the first count must be the number of workers, the second must be zero:
grep -h "Started database job worker" /tmp/worker-*.log | wc -l
grep -c "Started database job worker" /tmp/api.logStop everything with pkill -f "src.jobs.runtime"; pkill -f "server.py".
The lowered pool settings are not optional at this scale: PostgreSQL defaults to
max_connections=100, while 11 processes on the default pool of10 + 20can ask for up to 330 connections.
dev dependency group is used to distinguish from production ones.
For first-time local setup:
# install project dependencies (including dev group)
uv sync --dev
# NOTE: browser binaries are not installed automatically by `uv sync`
uv run playwright install# install production dependency
uv add mydep
# install dev dependency
uv add --dev mydepAll code should adhere to the quality checks below. It is highly recommended to instal pre-commit hook and also integrate these tools below in the IDE.
Quality checks using ruff and mypy.
uv run poe typecheck
uv run poe lint
uv run poe stylecheck
# optionally run all quality checks (including unit tests)
uv run poe qa
# attempt to fix formatting and lint errors
uv run poe fixRunning tests using pytest.
# run all tests
uv run poe test
# run in watch mode
uv run poe test-watch
# run unit tests only
uv run poe test test/unit
# run integration tests only
uv run poe test test/integrationThis project uses pre-commit to ensure consistent code style and other quality checks.
# install the hooks from .pre-commit-config.yaml
# this needs to be done just once when setting up project
uv run pre-commit installOnce installed, the hooks will automatically run every time you commit changes. If any issues are found or files are modified, the commit will be aborted until fixed.
For development and testing purposes is every api request and llm call traced with Langfuse. By default is langfuse tracing disabled, you can enable it by configuration:
# configure correct langfuse project keys
LANGFUSE__SECRET_KEY=project-secret-key
LANGFUSE__PUBLIC_KEY=project-public-key
# when using for development define your own environment
LANGFUSE__ENVIRONMENT=dev-myname
# enable langfuse
LANGFUSE__TRACING_ENABLED=true
- API endpoints and parameters have to follow camel case convention
- LLM model can be configured via any OpenAI-compatible chat API