Skip to content

Aggregator: deception and outlier detection beyond prompt hardening #21

Description

@stubbi

Add structural deception/outlier detection on top of the M1 prompt-level hardening, since a single deceptive member can otherwise nullify the gains (arXiv 2503.05856).

Acceptance: an adversarial test where one member returns a confidently wrong answer; the synthesis does not adopt it.

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:corechorus-core enginesecuritysecurity-sensitive

    Type

    No type

    Projects

    No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions