The Keystore Cleanup Service provides automatic deletion of cryptographic keys from the keystore for tokens that have been deleted (spent, expired, or invalidated). This ensures that the keystore doesn't accumulate stale keys indefinitely, improving security and reducing storage overhead.
The cleanup system consists of four main components:
- Manager: Orchestrates the cleanup process with periodic scanning and distributed coordination
- SKI Extraction System: Pluggable architecture for deriving Subject Key Identifiers (SKIs) from owner identities
- Storage Interface: Provides database operations for querying deleted tokens and tracking cleanup state
- Keystore Interface: Abstracts key deletion operations
The Manager runs in the background and periodically scans for deleted tokens that are eligible for key cleanup. It uses distributed locking (PostgreSQL advisory locks) to ensure only one replica in a multi-instance deployment performs cleanup at a time.
Key Features:
- Periodic scanning with configurable intervals
- Worker pool for parallel key deletion (created per sweep)
- Distributed leadership via advisory locks
- Idempotent operations with cleanup tracking
- Immediate initial sweep on startup
Lifecycle:
- Leadership is acquired per sweep, not held continuously
- Initial sweep runs immediately on start (before first interval)
- Worker pool is created and destroyed for each sweep
- Context cancellation stops workers gracefully mid-sweep
The SKI (Subject Key Identifier) extraction system uses a pluggable provider architecture to support different identity types. This allows identity-type-specific logic to be encapsulated in separate providers.
Components:
-
TypedSKIProvider Interface: Defines the contract for identity-type-specific SKI extraction
type TypedSKIProvider interface { // GetSKIsFromIdentity derives one or more SKIs from an identity's raw bytes GetSKIsFromIdentity(ctx context.Context, identity []byte) ([]string, error) }
-
SKIExtractor: Orchestrates SKI extraction by maintaining a registry of providers
- Maps identity types (e.g., "idemix", "x509") to their specific providers
- Routes extraction requests to the appropriate provider
- Falls back to default provider for unknown types
- Thread-safe for concurrent use after initialization
-
Built-in Providers:
- IdemixSKIProvider: Extracts SKI from Idemix NymPublicKey
- IdemixNymSKIProvider: Extracts SKI from Idemix pseudonym identities
- NoopSKIProvider: Derives no SKIs, marking an identity type as deliberately excluded from cleanup. Registered for X.509 — see X.509 is intentionally out of scope
- FallbackSKIProvider: Computes SHA256 hash of identity bytes as SKI (default)
Provider Registration:
extractor := NewSKIExtractor()
extractor.RegisterProvider("idemix", idemix.NewSKIProvider())
extractor.RegisterProvider("idemixnym", idemixnym.NewSKIProvider(identityStore))
// X.509 keys belong to the wallet, not to individual tokens: deliberately never cleaned up
extractor.RegisterProvider("x509", NewNoopSKIProvider())
// Fallback provider is used for any unregistered typesSKI Derivation Process:
- SKIExtractor receives identity bytes and type
- Looks up registered provider for that type
- If found, delegates to type-specific provider
- If not found, uses fallback provider (SHA256 hash)
- Returns list of SKI strings in hexadecimal format
X.509 is registered with NoopSKIProvider, so no keystore key is ever deleted for an
X.509-owned token. This is deliberate, not an unimplemented provider.
An X.509 owner identity is a long-lived, non-anonymous certificate: x509.KeyManager reports
Anonymous() == false and always serves the same identity descriptor. Its private key therefore
belongs to the wallet, not to any individual token — the same key signs every token that
wallet ever owns, and it remains in use long after those tokens are spent. A real X.509 SKI
provider would make the cleanup sweep delete the wallet's own signing key as soon as the first of
its tokens aged past the TTL, permanently breaking the wallet.
The Idemix providers are the opposite case: their SKIs identify a one-shot pseudonym key created for a single recipient identity. That key is dead once its token is deleted, which is exactly what makes it safe to remove.
Consequence to be aware of: a deleted X.509-owned token still gets a row in
token_ski_cleanups, because the cleanup manager records "no key material to delete" the same way
it records a completed deletion. For X.509 that row means nothing to clean, not keys were
removed. Do not read the token_ski_cleanups table as evidence that X.509 key material was
purged.
The Storage interface abstracts database operations needed for cleanup:
type Storage interface {
// AcquireCleanupLeadership obtains distributed lock for leader election
AcquireCleanupLeadership(ctx context.Context, lockID int64) (Leadership, bool, error)
// GetDeletedTokensPendingSKICleanup queries deleted, owned tokens older than TTL that haven't been cleaned.
// Tokens for which this node is only an auditor or issuer are excluded, since this node
// never holds the secret keys for those tokens.
GetDeletedTokensPendingSKICleanup(ctx context.Context, olderThan time.Duration, limit int) ([]DeletedToken, error)
// MarkTokenCleaned records successful key cleanup to prevent reprocessing
MarkTokenCleaned(ctx context.Context, txID string, index uint64, cleanedBy string) error
}
type Leadership interface {
// Close releases the leadership lock
Close() error
}
type DeletedToken struct {
TxID string
Index uint64
OwnerIdentity []byte
OwnerType string
DeletedAt time.Time
}type SKIProvider interface {
// GetSKIsFromIdentity derives one or more SKIs from an owner identity
GetSKIsFromIdentity(ctx context.Context, identity []byte, identityType string) ([]string, error)
}type Keystore interface {
// Delete removes the key with the given identifier
Delete(id string) error
// Close closes the keystore
Close() error
}
type KeystoreProvider interface {
// Keystore returns the keystore for the given TMS
Keystore(tmsID token.TMSID) (Keystore, error)
}Cleanup behavior is controlled via configuration (see Configuration):
services:
storage:
cleanup:
enabled: false # Disabled by default - must be explicitly enabled
ttl: 24h # Minimum age before cleanup
scanInterval: 1h # How often to scan
batchSize: 100 # Max tokens per scan
workerCount: 1 # Parallel workers (default: 1)
advisoryLockID: 0x74746b636c65616e # Lock ID for leader election ("ttkclean" in hex)
instanceID: "cleanup-1" # Instance identifier (auto-generated if empty)Configuration Details:
- enabled: Must be explicitly set to
trueto activate cleanup (conservative default) - ttl: Minimum age of deleted tokens before their keys are eligible for cleanup
- scanInterval: How frequently the manager scans for eligible tokens
- batchSize: Maximum number of tokens processed in a single sweep
- workerCount: Number of parallel workers processing tokens within a sweep (default: 1)
- advisoryLockID: PostgreSQL advisory lock ID for distributed coordination
- instanceID: Identifies the cleanup instance; auto-generated as
cleanup-<pointer>if not provided
Configuration Loading:
The configuration is loaded from the TMS configuration using the key services.storage.cleanup. The LoadConfig() function merges provided values with defaults, preserving defaults for any unset values.
import (
"github.com/LFDT-Panurus/panurus/token"
"github.com/LFDT-Panurus/panurus/token/services/storage/services/cleanup"
)
// Create configuration
config := cleanup.Config{
Enabled: true,
TTL: 24 * time.Hour,
ScanInterval: 1 * time.Hour,
BatchSize: 100,
WorkerCount: 4,
AdvisoryLockID: 0x74746b636c65616e,
InstanceID: "cleanup-instance-1",
}
// Create TMS ID
tmsID := token.TMSID{
Network: "fabric",
Channel: "mychannel",
Namespace: "token-chaincode",
}
// Create manager
manager := cleanup.NewManager(
logger, // logging.Logger
storage, // Storage interface implementation
skiProvider, // SKIProvider interface implementation
keystoreProvider, // KeystoreProvider interface implementation
tmsID, // Token Management System ID
config, // Configuration
)
// Start cleanup (runs initial sweep immediately)
if err := manager.Start(); err != nil {
return err
}
// Stop cleanup gracefully
defer manager.Stop()To add support for a custom identity type:
// Implement TypedSKIProvider interface
type CustomSKIProvider struct {
// your fields
}
func (p *CustomSKIProvider) GetSKIsFromIdentity(ctx context.Context, identity []byte) ([]string, error) {
// Parse identity and extract SKIs
// Return SKIs as hexadecimal strings
return []string{"ski1", "ski2"}, nil
}
// Register with extractor
extractor := cleanup.NewSKIExtractor()
extractor.RegisterProvider("custom-type", &CustomSKIProvider{})The cleanup service is typically integrated automatically through the service manager:
import (
"github.com/LFDT-Panurus/panurus/token/services/storage/services/cleanup"
)
// Create service manager (manages cleanup instances per TMS)
cleanupManager := cleanup.NewServiceManager(
configuration, // Configuration interface
identityStorageProvider, // Identity storage provider
tokensProvider, // Tokens service manager
)
// Service manager automatically:
// - Creates cleanup manager per TMS
// - Registers built-in SKI providers (idemix, idemixnym, x509)
// - Sets up storage and keystore adapters
// - Starts the cleanup manager- Startup: Manager starts and runs initial sweep immediately
- Leadership Acquisition: Manager attempts to acquire PostgreSQL advisory lock
- Query Eligible Tokens: If leadership acquired, query deleted tokens older than TTL that haven't been cleaned yet
- Create Worker Pool: Spawn configured number of workers for this sweep
- Distribute Work: Fan out tokens to worker pool via channel
- Per Token Processing:
- Get keystore for TMS
- Derive SKIs from owner identity using appropriate provider
- Delete each SKI from keystore
- Mark token as cleaned in database only if every key was deleted (or if the owner type has no keys to delete at all)
- Release Leadership: Close leadership lock
- Wait: Sleep until next scan interval
- Repeat: Go to step 2
Timing Details:
- Initial sweep runs immediately on
Start(), before first interval - Leadership is acquired and released for each sweep
- Worker pool is created per sweep, not persistent
- Context cancellation stops workers gracefully
The cleanup service handles errors gracefully with specific retry behavior:
Key deletion is all-or-nothing per token:
- Any key fails to delete: Token is NOT marked as cleaned; the whole token is retried on the
next sweep. The returned error joins every per-key cause, so callers can inspect them with
errors.Is/errors.As - All keys deleted: Token IS marked as cleaned
- No SKIs derived: Token IS marked as cleaned to avoid infinite retries; logs warning. This is the normal path for X.509 — see X.509 is intentionally out of scope
Marking a token cleaned while some of its key material is still in the keystore would turn a
transient database error into a permanent, un-retriable key-retention hole: the token would never
be selected again, so the surviving key would never be deleted. Retrying the whole token is safe
and cheap because Keystore.Delete is idempotent — re-deleting the keys that already succeeded
costs one no-op call each.
Empty SKI cases are still marked complete, otherwise every sweep would rescan the same tokens forever.
- Key Not Found: Logged as warning, continues with other keys
- Database Errors: Logged and retried on next scan
- Leadership Loss: Cleanup stops gracefully, another instance takes over
- Context Cancellation: Workers stop gracefully mid-sweep
A token transitions through the following states related to cleanup:
- Active: Token is unspent and in use
- Deleted: Token marked as deleted (
is_deleted=true,spent_atset) - Eligible for Cleanup: Deleted, owned token (
owner=true) older than TTL without a cleanup record intoken_ski_cleanups - Cleaned: Keys deleted from keystore (record exists in
token_ski_cleanups)
The cleanup service only processes tokens in the "Eligible for Cleanup" state. Tokens for which
this node is only an auditor or issuer (owner=false) are never selected, since this node
never holds the secret keys for tokens it does not own.
The cleanup service uses a dedicated tracking table to record cleanup operations:
Token SKI Cleanups Table:
CREATE TABLE IF NOT EXISTS token_ski_cleanups (
tx_id TEXT NOT NULL,
idx INT NOT NULL,
cleaned_at TIMESTAMP NOT NULL,
cleaned_by TEXT NOT NULL,
PRIMARY KEY (tx_id, idx),
FOREIGN KEY (tx_id, idx) REFERENCES tokens
);
CREATE INDEX IF NOT EXISTS idx_cleaned_at_token_ski_cleanups ON token_ski_cleanups ( cleaned_at );This table tracks when each token's cryptographic keys were cleaned from the keystore, preventing reprocessing. The cleaned_by field records which cleanup instance performed the operation, useful for debugging in multi-instance deployments.
The token_ski_cleanups table is automatically created by the schema initialization and does not require manual database alterations.
- Multi-Instance Support: Uses advisory locks for distributed coordination
- Leader Election: Only one replica performs cleanup sweeps at a time
- High Availability: Multiple replicas can share the same database
- Automatic Failover: If leader fails, another replica acquires leadership on next scan
- Lock Scope: Leadership is per-sweep, not held continuously
- Single-Node Only: SQLite lacks advisory lock mechanism
- Node Restart Support: Cleanup resumes automatically after restart
- Not Recommended: For multi-replica deployments, use PostgreSQL
- Enabled:
false(must be explicitly enabled) - TTL: 24 hours (ensures tokens are truly finalized)
- Scan Interval: 1 hour (less aggressive than recovery's 5 seconds)
- Batch Size: 100 tokens per sweep
- Worker Count: 1 parallel worker (conservative default)
- Advisory Lock ID:
0x74746b636c65616e("ttkclean" in hex) - Instance ID: Auto-generated as
cleanup-<pointer>if not provided
-
For High-Volume Environments:
- Increase
batchSizeto 200-500 for more tokens per sweep - Increase
workerCountto 4-16 for faster parallel processing - Decrease
scanIntervalto 30m for more frequent cleanup
- Increase
-
For Resource-Constrained Systems:
- Keep
workerCountat 1 to minimize CPU usage - Increase
scanIntervalto 2-4h to reduce database load - Keep default
batchSizeto limit memory usage
- Keep
-
For Security-Sensitive Deployments:
- Decrease
ttlto 12h for faster key removal - Decrease
scanIntervalto 30m for more frequent cleanup - Monitor cleanup metrics to ensure timely processing
- Decrease
-
For Multi-Instance Deployments:
- PostgreSQL Required: Multi-instance deployments require PostgreSQL for distributed coordination
- Keep default
advisoryLockIDunless running multiple independent cleanup systems - Consider setting explicit
instanceIDvalues for easier debugging and monitoring
- Each scan queries the token database, so
scanIntervaldirectly affects database load workerCountaffects CPU utilization during cleanup sweepsbatchSizeaffects memory usage and the duration of each cleanup sweep- The relationship
scanInterval < ttlensures timely cleanup without premature processing
Key metrics to monitor:
- Cleanup Rate: Tokens cleaned per hour
- Backlog Size: Number of tokens eligible for cleanup
- Error Rate: Failed cleanup attempts (check logs for details)
- Leadership Changes: Frequency of leader election (should be stable)
- Processing Time: Duration of each cleanup sweep
- Retried Tokens: Tokens that failed at least one key deletion and are still pending. A token stuck here across many sweeps means a key deletion is failing persistently, not transiently
Log Levels:
INFO: Successful cleanup operations, manager start/stopWARN: Failed key deletions, key not found, leadership issuesDEBUG: Detailed sweep information, SKI derivation, leadership acquisition
- TTL Safety: 24-hour default ensures tokens are finalized before key deletion
- Idempotency: Safe to retry cleanup operations
- Audit Trail:
token_ski_cleanupstable provides cleanup history with timestamps and instance tracking - Key Isolation: Only deletes keys for deleted tokens, never active tokens
- All-or-Nothing Marking: A token is only recorded as cleaned once all of its keys are gone, so a failed deletion can never be silently forgotten
- X.509 Exclusion: X.509 key material is never deleted by design — a
token_ski_cleanupsrow for an X.509-owned token means "nothing to clean". See X.509 is intentionally out of scope - Instance Tracking:
cleaned_byfield records which instance performed cleanup
| Feature | Recovery Service | Cleanup Service |
|---|---|---|
| Purpose | Re-register finality listeners | Delete stale cryptographic keys |
| Frequency | Every 5 seconds | Every 1 hour |
| TTL | 30 seconds | 24 hours |
| Target | Pending transactions | Deleted tokens |
| Urgency | High (affects finality) | Low (housekeeping) |
| Batch Size | 100 | 100 |
| Workers | 4 | 1 (default) |
| Initial Sweep | Immediate | Immediate |
| Leadership | Per sweep | Per sweep |
- Storage Service - Database operations and interfaces
- Configuration Guide - Detailed configuration parameters
- Transaction Recovery Service - Related recovery mechanism
- Identity Services - Identity management and SKI derivation