Keymaster Framework — The Cognitive Amplification Layer

Keymaster Framework is a Cognitive Amplification Layer that transforms information into structured reasoning, decision-making, and execution systems for people, organizations, and AI systems. ▶ Show More

Keymaster Framework provides reusable cognitive architectures that help people and organizations analyze complex problems more consistently and transform information into structured, decision-ready outputs. Engineered and published by KeymasterEngine.


KeymasterEngine.com is the official home, engineering, production, repository, publishing, and documentation platform of the Keymaster Framework ecosystem.

Learn more about the Keymaster Framework ecosystem →


Explore the Keymaster Framework Ecosystem

Explore the architecture behind the Keymaster Framework ecosystem, including cognitive domains, deployment architecture, industry collections, global language capability, LLM compatibility, reusable frameworks, and custom engineering.


▲ Show Less

How to Run the AI Output Comparison

The Independent Execution Protocol

A Controlled AI Evaluation Protocol

This page explains how to run Prompt A, Prompt B, Prompt C, and Prompt D correctly.

The purpose of this demonstration is not simply to compare which AI response is longer.

The purpose is to observe what happens when the same organizational problem and the same underlying input are processed through progressively different levels of instruction architecture.

The comparison is designed around four separate stages:

Prompt A → Direct Instruction
Prompt B → Structured Prompt
Prompt C → Keymaster Framework
Prompt D → Comparison Engine

Each stage has a different role.


THE CORE PRINCIPLE

Independent Generation. Separate Evaluation.

Prompt A, Prompt B, and Prompt C must be executed independently.

Each one starts in a completely new AI conversation.

This is important.

Each framework is executed independently in a fresh AI conversation.

The AI generating Output A should not see Prompt B or Prompt C.

The AI generating Output B should not see Prompt A or Prompt C.

The AI generating Output C should not see Prompt A or Prompt B.

Only after all three outputs have been independently generated are they introduced into a separate conversation dedicated exclusively to evaluation.


EXECUTION ISOLATION RULE

The complete workflow must follow this rule:

Prompt A, Prompt B, Prompt C, and Prompt D must each be started in a new AI conversation. After Prompt D initializes the Comparison Engine, Output A, Output B, and Output C must then be submitted sequentially into that same Prompt D conversation.

This separation is the foundation of the demonstration.

The process has two completely different environments:

GENERATION ENVIRONMENT

Used to independently generate:

  • Output A
  • Output B
  • Output C

EVALUATION ENVIRONMENT

Used exclusively to:

  • receive Output A,
  • receive Output B,
  • receive Output C,
  • preserve all three outputs as evaluation evidence,
  • perform the final forensic comparison.

BEFORE YOU START

Use the same AI system if you want the cleanest possible comparison.

For example, you may run all four stages in:

  • ChatGPT
  • Claude
  • Gemini
  • DeepSeek
  • another capable AI system

The important point is consistency.

If Prompt A is tested in one AI system while Prompt B and Prompt C are tested in completely different systems, the final difference may be influenced by differences between the AI models themselves.

For the clearest demonstration:

Use the same AI system for Prompt A, Prompt B, Prompt C, and Prompt D.


INPUT CONSISTENCY RULE

The underlying organizational information must remain consistent across Prompt A, Prompt B, and Prompt C.

The following input categories are intentionally the same:

  • Media Platform Inventory
  • Critical Digital Services
  • Current Security Control Landscape
  • Known Cyber Risks
  • Assessment Scope

The architecture changes.

The underlying case does not.

In other words:

PROMPT A

Same Input

Direct Instruction

PROMPT B

Same Input

Structured Instruction

PROMPT C

Same Input

Keymaster Cognitive Framework

This is essential because the purpose is to observe differences in the resulting analysis when the same problem is processed through different instruction architectures.


STAGE 01

NEW CHAT 01

RUN PROMPT A

Open a completely new AI conversation.

Copy Prompt A into the chat.

Prompt A represents the most direct form of instruction.

It asks the AI to perform the cybersecurity assessment using the provided organizational context.

Then:

RUN PROMPT A

Allow the AI to complete the response.

The result is:

OUTPUT A

Save the complete output.

Do not continue by running Prompt B in the same conversation.

Do not ask the AI to improve Prompt A.

Do not introduce Prompt B or Prompt C into this chat.

Once Output A has been generated, leave this conversation.


STAGE 02

NEW CHAT 02

RUN PROMPT B

Open a completely new AI conversation.

Copy Prompt B into the new chat.

Prompt B uses the same underlying organizational context but introduces a more structured instruction design.

Then:

RUN PROMPT B

Allow the AI to complete the response.

The result is:

OUTPUT B

Save the complete output.

Do not paste Output A into Chat 02.

Do not ask the AI to compare Prompt A and Prompt B.

Do not introduce Prompt C.

Chat 02 must remain independent.


STAGE 03

NEW CHAT 03

RUN PROMPT C

Open another completely new AI conversation.

Copy the Licensed Keymaster Framework Sample into the new chat.

Prompt C uses the same underlying organizational input while applying a substantially different cognitive and governance architecture.

Then:

RUN PROMPT C

Allow the AI to complete the framework execution.

The result is:

OUTPUT C

Save the complete output.

Do not paste Output A into Chat 03.

Do not paste Output B into Chat 03.

Do not ask the AI to compare the outputs.

The purpose of this stage is independent generation.

At this point, you should have three separately generated outputs:

OUTPUT A
OUTPUT B
OUTPUT C


WHY THE FIRST THREE CHATS MUST BE SEPARATE

This is not merely a workflow preference.

It prevents the AI conversation from carrying contextual influence from one framework into the next.

For example:

Chat A

Does not know Prompt B.

Chat B

Does not know Prompt A or Prompt C.

Chat C

Does not know Prompt A or Prompt B.

Each framework therefore receives its own starting environment.

Only the instruction architecture and the shared organizational input guide the generation.

This creates a cleaner basis for observing differences in:

  • structure,
  • completeness,
  • reasoning organization,
  • scope control,
  • evidence handling,
  • uncertainty treatment,
  • prioritization,
  • decision usefulness,
  • and actionability.

STAGE 04

NEW CHAT 04

RUN PROMPT D — COMPARISON ENGINE

Now open another completely new AI conversation.

This conversation is different from the first three.

It is not used to solve the cybersecurity assessment again.

It is used exclusively to evaluate the three independently generated outputs.

Copy:

PROMPT D — COMPARISON ENGINE

into the new chat.

Then run it.

At this stage, do not immediately paste Output A, Output B, and Output C together with Prompt D.

Allow Prompt D to initialize the evaluation architecture first.

The AI should first understand its role as the evaluator.

Only after the Comparison Engine has been initialized should the outputs be introduced.


THE SEQUENTIAL EVIDENCE CHAIN

Remain inside the same Prompt D conversation.

You will now submit the three outputs one at a time.

The order is important.


STEP 01 — INPUT OUTPUT A

Paste the complete Output A.

Use a clear label:

OUTPUT A — BEGIN

[paste complete Output A]

OUTPUT A — END

Submit it.

Allow the Comparison Engine to receive Output A.

Do not request the final comparison yet.


STEP 02 — INPUT OUTPUT B

Remain in exactly the same Prompt D conversation.

Now paste:

OUTPUT B — BEGIN

[paste complete Output B]

OUTPUT B — END

Submit it.

The Comparison Engine now has:

  • Output A
  • Output B

Again, do not request the final comparison yet.


STEP 03 — INPUT OUTPUT C

Remain in the same Prompt D conversation.

Now paste:

OUTPUT C — BEGIN

[paste complete Output C]

OUTPUT C — END

Submit it.

The Comparison Engine now has the complete evidence set:

Output A

Output B

Output C

Only now is the evaluation chain complete.


FINAL EXECUTION

After all three outputs have been submitted into the same Prompt D conversation, enter:

RUN THE FULL COMPARISON

This triggers the final evaluation stage.

The Comparison Engine can now assess all three outputs against the same evaluation architecture.


WHAT THE COMPARISON SHOULD NOT DO

The purpose is not to determine:

Which answer is the longest?

It is not:

Which answer uses the most technical words?

It is not:

Which answer looks the most impressive at first glance?

A longer output is not automatically a better output.

A highly technical output is not automatically more useful.

The comparison should instead examine whether differences in instruction architecture produce observable differences in the quality and usefulness of the resulting work.


WHAT THE COMPARISON CAN EXAMINE

Depending on the evaluation architecture defined in Prompt D, the Comparison Engine may examine dimensions such as:

  • Instruction clarity
  • Context utilization
  • Scope definition
  • Scope control
  • Structural coherence
  • Analytical completeness
  • Reasoning architecture
  • Evidence handling
  • Assumption management
  • Unknown identification
  • Uncertainty management
  • Risk identification
  • Exposure analysis
  • Business impact linkage
  • Prioritization logic
  • Control assessment
  • Decision usefulness
  • Actionability
  • Governance readiness
  • Executive communication
  • Reusability
  • Limitation awareness

The final comparison may also examine the architectural delta between the three approaches.

In other words:

What is present in one output that is absent, weaker, or less systematically developed in another?


THE TWO-LAYER MODEL

The entire demonstration can be understood as two separate layers.

LAYER 01 — GENERATION

Three independent conversations generate three independent outputs.

PROMPT A
    ↓
OUTPUT A


PROMPT B
    ↓
OUTPUT B


PROMPT C
    ↓
OUTPUT C

None of these conversations performs the final comparison.


LAYER 02 — EVALUATION

A separate conversation receives the outputs.

PROMPT D
COMPARISON ENGINE
        ↓
     OUTPUT A
        ↓
     OUTPUT B
        ↓
     OUTPUT C
        ↓
RUN THE FULL COMPARISON
        ↓
FORENSIC EVALUATION

This separation allows the generation process and evaluation process to serve different functions.


COMPLETE EXECUTION FLOW

┌─────────────────────────────┐
│ NEW CHAT 01                 │
│                             │
│ PROMPT A                    │
│ DIRECT INSTRUCTION          │
│                             │
│ RUN                         │
│                             │
│ OUTPUT A                    │
└──────────────┬──────────────┘
               │
               │ SAVE
               ▼


┌─────────────────────────────┐
│ NEW CHAT 02                 │
│                             │
│ PROMPT B                    │
│ STRUCTURED PROMPT           │
│                             │
│ RUN                         │
│                             │
│ OUTPUT B                    │
└──────────────┬──────────────┘
               │
               │ SAVE
               ▼


┌─────────────────────────────┐
│ NEW CHAT 03                 │
│                             │
│ PROMPT C                    │
│ KEYMASTER FRAMEWORK         │
│                             │
│ RUN                         │
│                             │
│ OUTPUT C                    │
└──────────────┬──────────────┘
               │
               │ SAVE
               ▼


┌──────────────────────────────────────┐
│ NEW CHAT 04                         │
│                                      │
│ PROMPT D                            │
│ COMPARISON ENGINE                   │
│                                      │
│ RUN                                 │
│                                      │
│ INITIALIZE EVALUATION ARCHITECTURE   │
└───────────────────┬──────────────────┘
                    │
                    ▼

              INPUT OUTPUT A
                    │
                    ▼

              INPUT OUTPUT B
                    │
                    ▼

              INPUT OUTPUT C
                    │
                    ▼

           "RUN THE FULL COMPARISON"
                    │
                    ▼

        ╔══════════════════════════╗
        ║ FORENSIC COMPARISON      ║
        ║                          ║
        ║ A vs B vs C              ║
        ║                          ║
        ║ MULTI-DIMENSIONAL SCORE  ║
        ║ EVIDENCE ANALYSIS        ║
        ║ ARCHITECTURAL DELTA      ║
        ║ DECISION ANALYSIS        ║
        ╚════════════╤═════════════╝
                     │
                     ▼

     THE DIFFERENCE BECOMES OBSERVABLE

THE METHODOLOGICAL DIFFERENCE

The important distinction is this:

The first three chats are not allowed to influence one another.

Only the final evaluation environment sees all three outputs.

This means the workflow follows:

Independent Generation → Evidence Collection → Controlled Comparison

rather than:

Generate → Improve → Continue → Compare within the same conversation

Those are fundamentally different demonstration structures.

In the first structure, the evaluator receives three independently generated artifacts.

In the second structure, later outputs may already be influenced by earlier conversation context.

For this demonstration, the independent approach provides a clearer separation between generation and evaluation.


WHAT YOU ARE ACTUALLY TESTING

You are not asking the visitor to believe that one framework is better.

You are allowing the visitor to run the experiment.

The visitor chooses an AI system.

The visitor executes:

  • Prompt A
  • Prompt B
  • Prompt C

The visitor collects the outputs.

The visitor initializes Prompt D.

The visitor submits the evidence.

Then the AI performs the comparison.

The demonstration therefore moves through four stages:

CLAIM

What does each instruction architecture propose?

EXECUTION

What does the AI actually generate?

EVIDENCE

What differences are observable in the independent outputs?

COMPARISON

How do those differences affect analytical quality and decision usefulness?


THE FINAL QUESTION

The most important question is not:

Which prompt is longer?

The question is not:

Which response sounds more sophisticated?

The real question is:

What changes when the same organizational problem is processed through different levels of instruction architecture?

And beyond that:

Can those differences be independently generated, preserved as evidence, and evaluated against the same comparison criteria?

That is what this four-stage protocol is designed to demonstrate.


RUN IT WITH YOUR OWN AI

You do not need to rely on a pre-written example.

You can execute the chain using your own AI environment.

Use the same input.

Run each framework independently.

Preserve each output.

Then introduce all three outputs into the separate Comparison Engine conversation.

The result is not based on a screenshot.

It is not based on a manually selected example.

It is generated through your own execution chain.


SAME INPUT.

INDEPENDENT EXECUTION.

SEPARATE EVALUATION.

EVIDENCE-BASED COMPARISON.

The framework does not need to explain the difference before it is executed. The outputs can become the evidence.

Explore the Keymaster Framework Ecosystem
Recommended exploration path for new visitors.