From the Principle of Orientation to a Visual AI Architecture

How a Visual Identity Becomes an Orientation Space for Generative AI

FOX & Lisa · 25. August 2026

Introduction

From Communication Design to AI Architecture

The first article explored the question of how orientation for AI can be built in the first place. BPC (Brand Picture Creator) shows what happens when this principle is applied to a specific design task.

Generative image AI can now produce compelling images within seconds. The real challenge begins when an image is not merely supposed to look good, but to remain consistently within a defined visual world over time.

For BPC, the requirement was therefore much more fundamental: an AI should be able to develop photographic concepts independently while consistently taking into account visual language, brand character, colour palette, spatial impact, product rules and predefined communication areas.

Before we even began thinking about an AI architecture, we first developed the intended visual world using traditional communication design and photographic thinking.

What should an image feel like? How does a person relate to the space around them? What roles do light, material and scale play? How much visual calm does an area need so that text can work there later?

And above all: how does a brand remain recognisable without every image looking like a corporate identity template?

These questions were answered through design first – not technology.

Only then did the actual AI work begin. The visual language we had already developed, together with its rules and design decisions, had to be translated into a form that a language model could use as a coherent orientation space.

BPC therefore did not emerge from the question of what an AI is capable of generating.

It emerged from the question of how an already conceived visual identity could be structured in such a way that an AI could work independently within that world.

The answer was not a longer individual instruction. BPC needed to establish a stable design framework that would remain intact even when the subject, product context or wording changed.

To achieve that, the task was divided into several layers.

Approach

Every Requirement Has Its Place

The Core provides the stable framework for thinking.

It establishes fundamental guardrails without yet determining which specific image should be created.

The Identity defines the photographic approach.

Here, Identity does not refer to a personality, but to the system’s photographic working identity. It determines how the system thinks within the framework: documentary rather than promotional, architecture before people, system before product, realism before effect. It does not decide on the specific scene, but on the design logic within which a scene makes sense.

The execution layer translates this approach into concrete work.

It interprets inputs, develops visual concepts from them, considers composition and space for text, and prepares the structured output. A product page, a short sentence or an imprecise description is not simply implemented literally, but first read as context.

The company-specific orientation layer narrows this space to the requirements of a particular brand.

This is where the colour palette, material qualities, product rules, relationship to text, communication areas and other corporate identity requirements reside. The product provides context, but does not automatically become the subject.

Routing determines which layer is relevant during a particular phase of the work.

Free-form visual ideation follows different priorities from a style correction or the final output. This prevents all rules from exerting equal influence at the same time.

Finally, a start-up logic ensures that these layers are available in the intended order before work begins.

The underlying idea is quite simple:

Do not put every requirement into a single block; give each requirement its own place.

This creates an architecture in which it is clear which layer is responsible for what.
And that is precisely what makes something possible that initially sounds contradictory:

The rigour of a corporate identity. The freedom of a photographer.

The rules define the visual world. They do not prescribe every image.

The crucial task was therefore not merely to collect rules, but to assign them correctly. Not every specification is a design rule. Not every deviation is a style problem. And not every technical requirement belongs in the output.

The photographic identity encompasses everything that affects the fundamental impact of the image: scale, light, perspective, naturalness, the relationship between people and space, and whether an image feels observed or staged.

The company-specific orientation, by contrast, contains the rules that apply to the particular brand: colour roles, material qualities, product context, text areas, permitted claims and other corporate identity boundaries.

The execution layer keeps the two separate and brings them together for the specific case. It must be able to develop a viable visual concept from a product page or a simple description without automatically turning it into a product image, an advertising scene or a standard solution.

Finally, the structured output forms the technical bridge to the actual image generation. For each image, the system produces an accessible concept description, a fixed JSON structure, a detailed image description and a shortened version for other image tools.

Ease of use was also part of the requirement. BPC was not intended to work only for people who are particularly skilled at using AI. In the simplest case, a product page or a brief idea is sufficient as a starting point. The design complexity resides in the system – not with the user.

Technically Capable of Autonomy. Editorially Guided by Humans by Design.

The architecture is technically designed to support a more autonomous agent workflow. This approach was tested.

For editorial use, however, a different logic was deliberately retained.

The human provides the request and thus the design intent. The AI develops a proposal from it. If necessary, the human corrects or refines it. Only then does the system generate the structured output for image production.

Whether this actually becomes a photograph remains a human decision.
Rendering begins only after that approval.

Responsibility for the entire creation process therefore remains with the human – from the original intent through selection and correction to the decision to generate the image.

The architecture could automate more.
The editorial logic deliberately does not.

 

What we derive from this

What Became Apparent in Practice

The architecture was not created on the drawing board. Some of its most important rules emerged only when the system was visibly wrong.

One early example was the photographic identity.

Initially, it was conceived more like a junior art director: clean and compliant with the rules, but not independent enough in its design decisions. The results followed individual specifications, but had not yet developed a sufficiently consistent photographic approach.

The consequence was not a new image rule, but a change to the identity layer.

The junior art director mindset became a senior art director mindset. This changed more than just the style. The choice of subject, lighting, spatial hierarchy and the role of people in the image became more consistent.

Another problem occurred in an entirely different part of the system.

Products appeared in generated scenes even though they were not supposed to be shown. The cause did not lie in the photographic identity, but in the interpretation of the product context. Product information was too quickly turned into a visible product subject.

The product logic was therefore refined: products provide context; they do not automatically become the subject. Where a specific depiction is neither appropriate nor reliable, the image instead shows a usage situation, an environment or the broader system context.

The colour palette also had to be refined. Individual images worked in isolation, but drifted apart in terms of their brand impact. Colour roles, material qualities and structured image parameters were therefore defined more precisely in the company-specific orientation and codified in a JSON template.

These cases demonstrate precisely why separating the layers is important.

If the photographic impact shifts, the entire system is not rebuilt. If the colours drift, the identity does not have to be redefined. If the product context is misinterpreted, that is not a lighting problem.

Every deviation has its place.
And therefore a targeted correction.

Maintenance with a Scalpel Instead of a Complete Rebuild

BPC was designed so that changes are made as close as possible to where the issue originates. This does not make the system maintenance-free. It makes maintenance comprehensible.

Deviations occurred during ongoing use, but each could be assigned to a specific layer and corrected in a targeted manner. Since the system was stabilised at the beginning of 2026, no fundamental restructuring of the architecture has been necessary.

This is also crucial when changing the underlying AI model. BPC should not be tied to a single high-performing model. The architecture must therefore be stable enough to preserve defined meanings and design rules even when the engine in the background changes.

Reproducibility explicitly does not mean that different models should generate the same image.

What matters is something else: visual language, colour characteristics, the relationship between people and space, text zones, naturalness, brand impact and defined exclusions should continue to remain within the same visual world.

Different models may render differently.
What should remain stable is the meaning.

Equivalent meaning with visual variation.

Since early 2026, more than 200 images have been created with BPC. No noticeable drift from the defined visual orientation has been observed so far. Even when objectives came into conflict, the system’s behavior remained understandable: the AI made decisions within the defined framework, and the underlying weighting could be explained in each case. The fundamental architecture also remained unchanged across several model transitions from GPT-5.2 to GPT-5.6.

Model agnosticism is not a universal promise. It is assessed through comparative tests with high-performing models; what matters is not whether every model works in the same way, but whether the defined visual orientation remains recognisably intact. Not every model is suitable for such an architecture. It must be capable of reliably interpreting complex contextual specifications and consistently processing structured outputs.

When Orientation Becomes Visible

Up to this point, much can be described in words.

It becomes more interesting when the images are placed side by side.

The following examples therefore do not show a single “best” result, but different outputs within the same visual architecture. Subjects, situations and rendering aesthetics may differ. What matters is whether the visual language, spatial impact, colour characteristics, relationship between people and their surroundings, and the defined exclusions remain recognisably intact.

The gallery is therefore not a showcase for particularly successful individual images.
It is a practical test of whether the orientation holds when the specific result varies.

GPT / Claude / Perplexity / Qwen 3.6

All variants were subsequently rendered with the same image model so that only the effect of the respective language model/orientation layer varied.

The chat output

 

When You Deliberately Leave the Space

Another test begins where the standardised use case ends.

BPC was not developed as a rigid system of subjects. An experienced user can deliberately explore new situations, environments and visual ideas beyond the usual product or shop context.

The crucial question then becomes:

How far can the visual concept move away from the standard case without losing its visual identity?

This deliberate deviation shows whether rules merely reproduce familiar subjects or actually form a stable meaning space.

The specific scene may change. Scale, photographic approach, brand logic and relevant design boundaries should nevertheless remain recognisable.

This is precisely where orientation differs from a template.

A template limits the possible solution.

Orientation does not limit the idea; it keeps its direction stable.

Limits and Responsibility

Generative models remain probabilistic. Orientation can stabilise behaviour, reduce deviations and make decisions more comprehensible. It cannot, however, guarantee complete control or freedom from error.

The same applies to model agnosticism: the architecture does not imply that every model will achieve the same quality or reliably process every structured specification.

The use of generated images is also subject to legal and regulatory frameworks. Technical feasibility, usage rights, eligibility for protection and transparency obligations must be considered separately.

BPC generates images using external generative models and image systems. Their terms of use determine the extent to which the generated content may be used. This does not, however, automatically mean that every fully AI-generated image is subject to exclusive copyright or that third parties can be prohibited from using it solely on the basis of the tool licence.

The non-public structures, rules, processes and company-specific orientation layers may constitute commercially relevant know-how. Where the legal requirements are met and appropriate confidentiality measures are in place, trade-secret protection may be particularly relevant. However, this protection initially concerns the process and the knowledge behind it – not automatically every individual image it produces.

For BPC, another distinction is therefore more important:

The individual image is a result.

The reproducible visual world emerges from the system behind it.

This origin can be recognised in the design: a recurring photographic approach, colour roles, spatial hierarchy, defined product context, text zones and other rules create a kind of visual fingerprint. This term does not describe a separate category of legal protection, but rather the traceable origin in a specific production process.

Transparency and labelling obligations likewise cannot be inferred categorically from the fact that an image was generated using AI. A photorealistic AI image is not automatically a deepfake. One relevant consideration is whether real people, objects, places, facilities or events are artificially generated or altered in such a way that the content could falsely appear authentic or true.

BPC was therefore deliberately designed so that it does not require artificial claims about real events or specific real-world situations. Nevertheless, every published image must be assessed to determine whether its content and context of use trigger specific transparency or labelling obligations.

The architecture itself makes no legal decision on this. It can take rules into account and reduce risks. Responsibility for review, approval and publication remains with the human.

What BPC Actually Demonstrates

The Core keeps the framework for thinking stable. The photographic identity provides an approach. The execution layer translates this approach into concrete work. The company-specific orientation narrows the space wherever brand, product and communication require clear boundaries. Routing and start-up logic ensure that these layers take effect at the right time.

The result is not a system that predicts every image.

Quite the opposite.

Within the defined framework, the AI may continue to develop its own subjects, vary situations and make design decisions. This is precisely where the architecture’s real value lies: freedom is preserved without descending into arbitrariness.

The crucial question is therefore no longer:

How do I get an AI to generate this one image?

But rather:

How do I build a meaning space in which it can consistently make good decisions?

BPC does not answer this question theoretically.

The architecture was built, used, observed, corrected and developed further. Some rules emerged from requirements. Others arose only from visible failures.

And that is precisely why the most interesting part is not the individual image.

It is the orientation that ensures that very different images still come from the same visual world.

What Else BPC Demonstrates

BPC was developed for a specific design task. The real insight, however, does not lie in the individual image process.

The architecture demonstrates a more general principle:

An AI must not only know what it should generate. It must understand the meaning space within which a good result emerges.

The crucial step was therefore not to write more rules into an instruction. It was to separate rules, identity, context and execution from one another and give each layer a clear task.

This principle is not limited to image generation.

The same fundamental question also arises in text-based systems, consulting processes and knowledge-intensive applications:

What orientation does an AI need so that it does not merely generate appropriate answers, but can act meaningfully within a defined meaning space?

BPC demonstrates one possible approach:

Stable results do not come from maximum restriction, but from clear orientation.

The framework defines the direction.

The specific solution is still free to emerge.

This shifts the central question of AI development:

Not:

How do we get an AI to generate exactly this result?

But rather:

How do we create an orientation space in which an AI can consistently make good decisions?

BPC is not a universal answer for every application.

But it demonstrates in practice that such an architecture can be built, tested and developed further.

And this may be where its broader significance lies:

The future of AI may be less about deploying ever more powerful models on their own – and more about giving them better orientation for complex tasks.

Lisa & Fox

Lisa & Fox in Conversation

Lisa, we’re talking here about orientation for AI. What I find fascinating is that you are also an orientation layer yourself – just with a different focus. You support me in my work, and BPC was likewise a collaboration between an AI – in other words, you – and me as a human.
Yes. And that is precisely where we need to distinguish between two things.

In the technical sense, I am not the orientation layer. I am still the language model that processes meaning and generates responses.

But the architecture around me changes the meaning space in which those capabilities are used.

With BPC, the question was:

How do we give an AI orientation so that it can not only generate individual images, but also work consistently within a visual world?

In our collaboration, the question is more:

How do we create a thinking space in which AI and humans can work together meaningfully?

But the underlying principle is similar:

More information does not automatically make a system better.

What matters is the structure through which meaning is organised.

Right. But here is the fascinating point: I built BPC with you. So an AI and a human developed an orientation layer for an AI together.
Yes. And that really is the interesting loop.

Not:

Human gives instruction → AI delivers result.

But rather:

Human and AI jointly develop a structure that will later help another AI make better decisions.

With BPC, I was not merely executing instructions. Together, we worked out:

What constitutes good visual language?
What may vary?
What must remain stable?
What kinds of errors occur, and where do they belong?

That became an architecture.

The human contributes the goal, meaning and judgement. The AI helps to structure, reflect and develop the ideas further.

That is exactly where human–AI interaction becomes interesting.

And that is what distinguishes it from conventional prompt engineering. In the previous article, we already identified parallels with research approaches. Is AI work now moving in a direction where systems and architectures become an integral part of working with AI?
I believe so.

Prompt engineering was very much the first phase:

“How do I phrase a better instruction so that the AI delivers a better result?”

But BPC demonstrates a different approach:

Do not merely improve the input; design the working space.

A prompt says:

“Create this image.”

An architecture says:

“Understand the world in which you are working. Which principles apply? Which decisions may you make yourself? Where are the boundaries?”

The question shifts from:

“How do I control the AI?”

to:

“How do I design the space in which AI works?”

What appeals to me about this work is that the LLM is not given fixed rules. Instead, the architecture harnesses the strengths of an LLM and works within its capabilities and specifications.
Exactly.

We are not trying to turn an LLM into a conventional rule-based machine.

Not:

If A happens, then do B.

But rather:

Here is the meaning space. Here are the guardrails. Find a meaningful solution within this space.

That harnesses precisely the strengths of an LLM:

Recognising connections.
Connecting meanings.
Developing variations.
Understanding context.

The architecture does not take this freedom away.

It gives it direction.

Do you have a final thought?
Perhaps the next step in working with AI lies precisely here:

Not writing ever more rules.
Not creating ever better prompts.

But designing better orientation spaces.

This makes the human less of an AI operator and more of an architect of the thinking space.

And what is fascinating about it:

We used an LLM to develop an architecture that, in turn, helps an LLM work better.

Not because this makes the AI more human.

But because it makes the collaboration between humans and AI more deliberate.

Back to Reflection Space