Essential cookies are active for security and service continuity. You can also enable optional analytics cookies.
Cookie summary:
Disclaimer summary: This is a research product/service offered without warranty, guarantee, or promise regarding availability, accuracy, performance, security, or any other aspect.

AI-native Delivery Experience Framework - IntDEx_v4.md
<!-- AI-native Delivery Experience - IntDEx. Provided under CC BY 4.0. https://IntDEx.org --><!--WHAT THIS FILE ISThe IntDEx methodology reference. It explains WHY the engine behaves asit does. The other seeds tell the model what to do; this one is thereasoning behind those rules, and is what a human reads to understand or challenge them.WHEN IT IS READReferenced per message, consulted when a rule needs interpretation.TO REUSE THIS FILE ON A NEW PROJECTCopy it VERBATIM. It is the framework itself, not project configuration.YOU CAN EDIT THE FRAMEWORK TO TAILOR IT TO YOUR ORGANISATIONIn this case, we suggest that you create an AI-Native Delivery Experience Framework (IntDEx) - {OrganisationName}.md file and reference it in this document.Within the AI-Native Delivery Experience Framework (IntDEx) - {OrganisationName}.md file, include a reference to IntDEx_v{v}.md. This approach allows you to clearly identify the differences between the original framework and your customised version, while maintaining and updating them independently. This also simplifies the adoption of future updates to the original framework.-----------------------------------------------------------------------------CONTENT FOR A POSSIBLE ORGANISATION'S CUSTOM FRAMEWORK FILE:CHANGE LOG:Date of Change:User:Description:Date of Change:User:Description:Ling to original IntDEx Framework: .pdm\AI-native Delivery Experience Framework - IntDEx_v2.md=============================================================================--># AI-native Delivery Experience - IntDEx (v2)## Link to the Organisation's Tailored Framework-None## Artefact Metadata- Created: 2026-09-06 22:33:20 +02:00- Created By: engine- Stage: engine- Work Item: WI-INTDEX-V2-DOCX-IMPORT- Timestamp Source: engine-clock> Source: `AI Native Delivery Experience Framework - IntDEx_Pub2.docx` (converted verbatim; headings, lists and tables mapped to Markdown)._An engineering framework for AI-native development and management of digital solutions.__Spyros Ktenas,_ _September_ _2026_## Definitions introduced in this paper**IntDEx:** Stands for (Artificial) Intelligence-native Delivery Experience.**AI-native delivery:** Is used to identify the complete (from inception to retirement) project/product delivery flow where Artificial Intelligence (AI) is the foundational operating/tooling layer. Part of AI-native delivery is the AI-native development that in IntDEx is the technical (aanalysis, coding, configuration) part of the digital solution. It covers the whole digital solution engineering. In AI-native delivery/development, removing AI wouldn’t just slow down the flow, it will break it.**Generative** **creation**: Is the AI-native build phase in IntDEx, where prompts, models, constraints, and context produce behavioral artefacts. It shifts emphasis from traditional coding and configuration with AI generation and refinement.## What IntDEx brings as new to existing practicesWhile many of IntDEx practices are close to existing spec-driven frameworks, e.g. the notion that an executable specification, rather than source code, is the durable engineering asset is already present in behavior-driven development, and spec-driven development, IntDEx, brings a number of new elements that gap current issues.Issues that IntDEx addresses:- New bottlenecks as coding speed has been dramatically increased.- Cost management risk- Terminology misalignment- Unclear roles and responsibilities- Hallucinations (Schmitt, 2024)- Unvalidated code generation- Non-deterministic behavior (Shah, 2026)- Amplifying bad practices- Unenforceable governance- Unenforceable human controls- Unenforceable ethical guidelinesIntDEx defines a delivery model where AI is the main tool for all roles in the SDLC. With evidential independence, reproducibility from artefacts as evidential claim, enforceable human in the loop (HITL) checkpoints, traceable artefact mappings, consideration of ethical constraints, and cost accounting. In addition, it includes definitions and context descriptions to ensure common language.Note: Implementation guides and examples that have been validated in practice together with configuration files, metrics, templates etc. have been developed, however, these are not part of the present research paper.## AbstractArtificial intelligence has introduced a new paradigm in software engineering, where the traditional digital solutions (or software) development lifecycles (SDLCs) assuming relatively long-linear development stages should be changed for AI-native development, with non-deterministic methods that require continuous adaptation, refinement, and behavior-centric engineering with fast progress between different product development phases. This paper introduces IntDEx ((Artificial) Intelligence-native Delivery Experience), an AI-native engineering framework designed to support the delivery of digital solutions in complex environments using AI as the main enablement tool. The framework was refined through practical application in a research project where it was used to accelerate the development of the digital platform that was needed for the required research experiments. IntDEx define the basis for a complete AI-native product/project SDLC, including roles, governance, delivery flow phases, and behavioral engineering practices. It emphasizes non-linearity, behavior validation, intent-centric design, context management, model adaptivity and human centricity as its foundational characteristics. The paper positions IntDEx as a new as a whole -while some parts are coming from existing best practices and are presented here for completeness-, academically supported and practically tested methodology for AI-native digital solution development. Additional work could further develop and detail specific areas of the framework.## IntroductionThe development of large language models (LLMs) and AI-native engineering practices can (and already has in many industries and organisations) fundamentally transform how digital solutions are conceived, designed, and delivered. Unlike traditional software development, which rely on human generated deterministic content, AI-native methods exhibit fast non-deterministic, emergent behavior driven by intents and interactions between prompts, models, and data. This shift challenges the assumptions of current project/product management and development methodologies, which depend on linear progression from requirements to design, implementation, and testing. In addition, the increased speed of development can significantly change the current development workflows and turn timed -boxed stages practical usefulness. New bottlenecks may appear in the flow and finally, just sending messages to an AI agent to produce code (vibe coding) is nowhere close to safe for a production system (Graziano, 2024).IntDEx has considered CPMAI (PMI, Cognilytica, 2025), the PMI AI Standard (PMI, 2026). It also incorporates current AI-native engineering research and recommendations, including the AI-Native Engineering practices (Graziano, 2024), AI-Native Engineering Team guidance (OpenAI Developers, n.d.), the AI-Native Engineer Roadmap (ByteByteGo, 2026) and more. These sources collectively highlight the need for intent-centric design, behavior-driven engineering, model adaptivity, and continuous intelligence loops and suggest new practices and adaptation to existing development methodologies.IntDEx as an AI-native delivery framework proposes a production ready flow, roles and responsibilities and practices for the development of complex digital solutions. The framework, as it is presented in this paper, constitutes just an initial baseline, that has been development in parallel with experimentation and practical application. The development is ongoing and what is today in this paper will change during future framework adjustments and refinement. This paper has detailed different aspects at different levels of detail based on what was needed mostly for the practical application in specific delivery project, therefore, further work may be needed to detail specific areas of the framework.## MethodologyThe development of this paper followed a continual adjustment, parallel methodology in which the IntDEx framework, the literature review, and the practical implementation evolved simultaneously. Rather than applying a linear process, the framework was shaped through continuous cycles of refinement driven by real delivery needs in the associated research project. As a digital platform was being engineered, new behavioral patterns, gaps, and requirements emerged, prompting adjustments to the framework and updates to the paper’s structure and content. In parallel, relevant academic and industry sources were reviewed and integrated, ensuring that each iteration of IntDEx aligned with recent AI-native engineering practices. This co-evolution of theory and practice produced the first baseline version of IntDEx presented in this paper, with the recognition that further iterations will be required as the framework continues to mature.## Background and Foundational Sources### 5.1 Origins of IntDExThe IntDEx framework emerged from the practical needs of the research initiative _“Interoperability and utilisation of epidemiological surveillance data with artificial intelligence.”_ This project required rapid development of digital tools and analytical components, continuous model evaluation, and adaptive refinement of digital tools. Traditional software development methodologies were considered insufficient because of the time to implement, and knowledge gaps in the research team. To accelerate experimentation, data exploration, and model assessment, the project adopted AI-native development practices.During early implementation, it became clear that a structured, repeatable framework tailored specifically to AI-native development was required, however, literature review revealed no concrete framework available. IntDEx was created to fill this gap. The framework is designed to support complex environments where governance and control are important but, also all cases where rapid prototyping, continuous intelligence loops, and human-in-the-loop oversight are essential in order to deliver better and faster.The term IntDEx (Intelligence-native Delivery Experience) is introduced here as a novel methodological construct. While “Developer Experience” (DX) has been already introduced in industry literature at least from 2011 (Cohick, 2011), no existing framework defines DX in the context of AI-native development. IntDEx extends the concept by defining an engineering framework and methodology elements where AI is the primary interface and tool for delivering digital solutions.### 5.2 IntDEx and Hybrid Project/Product ManagementIntDEx suggests being used in a Hybrid Project/Product Management context where Agile practices can be applied in a plan driven, longer initiative. This hybrid approach is described in multiple well-established methodologies. SAFe’s Big Picture (Worldofagile.com, n.d.) contributes flow-based delivery concepts, value stream alignment, and governance structures suitable for large-scale initiatives in this hybrid mode. Ktenas’ Software Project Management presentation (Ktenas, 2016) provides advice on how a hybrid approach could work. PRINCE2 methodology can include agile cycles in its “Product Delivery” stage and PM²-Agile Guide 3.0.1 (European Commission, 2021) provide details for integrating Agile practices into plan-driven environments, emphasising transparency, collaboration, and iterative delivery.IntDEx do not aim to replace Project/Product management methodologies but to provide a new framework that can easily be followed and implemented within established Project/Product Management contexts so that development teams can benefit the AI evolution. However, some adjustments to existing project product management frameworks and tailoring may be needed.**Flow-based delivery** instead of time boxed iterations or sprints**.** IntDEx does consider sprints and timeboxed iterations as a necessity. These could also remove the need for most of the events that are proposed by the different agile methodologies relevant to sprints/iterations. A kanban based approach is closer to the IntDEx approach. IntDEx propose flow-based delivery, a readiness-based replenishment rather than fixed iterations. This aligns with more articles advocating continuous flow over time-boxed iterations in complex environments (Denis Dennehy, 2018). However, IntDEx may coexist with sprint/iteration-based governance when required.**More and smaller teams** vs less and bigger teams**.** Team size should be guided mainly by the complexity of the business operations and needs, the business stakeholders map, regulatory implications, and business uncertainties and not that much of the anticipated technical implementation complexity. This is a major shift. The Delivery Manager could also be in position to estimate the effort/cost together with the Developer, or a technical expert. The IntDEx team should be dedicated as much as possible to a business domain rather than technology. Team size, however, cannot be less than three (3) people for continued delivery with an ideal number considered four (4) or five (5).### 5.3 AI-Native Engineering Research FoundationsIntDEx has considered recent AI-native engineering research. The requirements for Continuous orchestration, model adaptivity and human-AI collaboration noted in various articles (Charbel Daoud, 2026) . The importance of governable and transparency (OECD, n.d.). New practices for _AI-Native Software Development Lifecycle_ (Claxton, 2026). OpenAI’s _Building an AI-Native Engineering Team_ (OpenAI Developers, n.d.) outlines practical guidance for structuring engineering teams around AI-native workflows, including prompt versioning, model routing, and behavior validation. Finally, risk management is also different for AI (NIST, 2026).ByteByteGo’s _AI-Native Engineer Roadmap_ (ByteByteGo, 2026) highlights the skills and practices required for AI-native engineering, including prompt architecture, model evaluation, and continuous monitoring. Microsoft’s _Spec-Driven Development_ (Gupta, 2026) introduces a spec-first approach to AI-native engineering, where prompts act as executable specifications. InfoQ’s _Four Patterns of AI-Native Development_ (InfoQ, 2026) identifies architectural patterns such as retrieval-augmented generation (RAG), agent-based systems, and model-routing architectures.Additional pre-exiting relevant work has been considered (as referenced) in this document.### 5.4 Practical Research and ExperimentationIntDEx was refined through hands-on experimentation in a software development project. The research project required rapid development and prototyping, continuous model evaluation, and adaptive refinement of digital tools. This practical context validated the need for:- Nonlinear, flow-based delivery cycles- Emergence, needs and changes arise from many small interactions- Continuous behavior validation- Intent-centric development- Context and constrain management. AI systems respond differently in different contexts.- Adaptive model selection- Cost management- Provisions for documentation and supporting artefacts creation- Human augmentation. Expanding human intelligence, human-in-the-loop oversight, ethical safeguards, continuous monitoring for drift and hallucinationsTogether, pre-existing work and the practical application observations shaped IntDEx’s characteristics, delivery flow, roles, and practices.#### Details on Practical ValidationThe proposed approach has been used for the development of a data science platform and the relevant support tools in order to provide a standalone sovereign solution for public health data experimentation.During the project the AI-Native approach has been used, and an initial set of AI-Native practices, lifecycle phases and templates have been developed. For the baseline version of the solution GPT-5.3-Codex was used. Auxiliary AI services were also used (e.g. MS copilot). The IDE was Visual Studio Code. The IntDEx templates have been used to rebuild the solution from artefacts. This was successful as limited behavior deviation observed. For the behavior deviation calculation code comparison has been produced together with human UI tests. AI comparison was also used with different family models. Details on the resultartefact-regenerated reimplementation. Route surface ~95% identicalDeviations beyond business features noticed at- Architecture / config: different file names, structure- Data / migration: different db tables, columns etc.- Security regressions: weaker security has been implementedThe test was on custom codebase footprint of 52,598 LOC.With further configuration of the Delivery Engine the deviations could be further reduced.In addition, the specific engine configuration has been saved and will be used to continue the work presented here with examples and further evidence of the feasibility of IntDex practical implementationReproducibility from artefacts, documentation, governance, cost control, risk management, not speed, is the primary evidential claim of this framework. The solution was re-created only by the artefacts created by the IntDEx compliant AI-native delivery without the use for the codebase. This is important as it validates the engine and its artefacts as an accurate set of deliverables that can be used to validate, describe and explain the delivered behavior and demonstrate compliance with organisational/project constraints, guidelines and contents.## 6 IntDEx CharacteristicsAI-native development differs from traditional software engineering. Traditional SDLCs work better with human provided deterministic content, stable requirements (even in agile, requirements should not change during an iteration), and linear progression. IntDEx assumes that requirements can change as fast as they can be implemented and incorporates a number of defining characteristics that form the basis of the IntDEx framework.### 6.1 Non-LinearAI-native delivery is non-linear because engineering activities do not occur in a predetermined order. In traditional development (even in an agile context), work is arranged into sequential stages. AI-native delivery breaks this sequencing as intents, prompts, models, constraints, data, and human oversight (ISO, 2022) influence one another continuously, and changes in any of these elements can reshape the others immediately.This creates an environment defined by continuous interaction rather than staged progression. The AI delivery engine may change in response to model performance, drift, cost, latency, and safety considerations, while the digital solution changes when the delivered product or service behavior must be improved. Prompts evolve as versioned artefacts that respond to behavioral feedback. Delivery readiness is determined by behavioral quality rather than by iteration boundaries, reflecting principles of flow-based delivery. Monitoring provides real-time insight into delivery-engine drift, hallucinations, constraint degradation, and data quality, ensuring that engineering decisions remain responsive to the AI delivery engine’s evolving behavior.Non-linearity in IntDEx therefore describes the reciprocal influence of engineering concerns. AI-native systems require continuous, bidirectional adjustment across requirements, prompts, models, architecture, and oversight, making fixed sequences impractical and counterproductive.This means that requirements, implementation, testing, and architecture influence each other simultaneously rather than sequentially.IntDEx replace linear SDLC progression with a continuous intelligence loop, where any phase can trigger a return to any other phase. Requirements are expressed as intents rather than fixed user stories, and they evolve during development as behavior emerges. Pair interaction, collaboration between developers and analysts or testers, supports that requirements, implementation, and testing occur simultaneously.The framework is aligning with flow-based delivery principles advocated in SAFe (Worldofagile.com, n.d.) and flow research (Denis Dennehy, 2018).### 6.2 EmergentThe AI delivery engine can exhibit emergent behavior that arises from the interaction of prompts, models, constraints, data, context, and human oversight. The resulting digital solution may also exhibit emergent user-facing behavior, but it should be evaluated separately as the product or service delivered to users. This behavior cannot be fully pre-determined as it unfolds as the engine operates and as new information is introduced. Emergence reshapes expectations, architectural choices, validation strategies, and risk posture, because the system reveals capabilities and limitations that were not visible at design time.In IntDEx, emergence is treated as a defining property rather than an anomaly. Behavioral intent evolves as stakeholders refining their understanding of what the digital solution should do. Architectural components of the AI delivery engine, such as retrieval pipelines, safeguards, and model routing, adapt to new behavioral patterns. Prompts mature through iterative refinement and versioning. Validation exposes new constraints and edge cases. Risk posture shifts as the AI delivery engine encounters new contexts. Human oversight identifies behaviors that automated mechanisms cannot detect. Emergence therefore becomes a continuous source of insight that guides engineering decisions across the entire delivery effort.Emergence is not a new observation. Complex adaptive systems theory, the Cynefin framework, emergent design in agile practice, and emergent behavior in distributed systems all describe the same phenomenon, and IntDEx claims no discovery here. What differs is where emergence happens and the scale and speed or evolution. In traditional development, the work is deterministic and emergence appears at the level of the system, the organisation, or the requirement set (the code does what it says). In AI-native delivery the unit of work itself is non-deterministic.### 6.3 Behavior-DrivenIn AI-native delivery, the digital solution is defined by behavior, the observable output produced by prompts, models, constraints, data, context, and human oversight. Behavior becomes the primary delivered artefact, replacing code as the most important engineering object. Code and infrastructure will be managed by the AI delivery engine.IntDEx focuses on producing behavior rather than developing code. Validation examines behavioral correctness, consistency, and reliability. Integration ensures that behavior remains coherent across systems. Release packaging delivers behavioral artefacts. Monitoring tracks behavioral drift, bias, and hallucinations. Maintenance becomes the evolution of behavior rather than the modification of code.### 6.4 Intent-CentricIntDEx is an intent-centric framework. Intent becomes the primary AI drafted, human-validated specification, replacing traditional user stories, requirement documents, and functional specifications. Behavioral intent expresses what the system should do, how it should behave, which domain rules apply, which constraints must be respected, and which outcomes are acceptable.Intent is written in domain-specific language rather than technical syntax. It captures stakeholder expectations, subject-matter expertise, and governance requirements. Intent serves as a stable reference point against which behavior is evaluated. Prompts operationalise intent into executable instructions. Behavioral artefacts are generated with intent as their reference. Monitoring identifies deviation between actual behavior and intended behavior. As domain rules or expectations change, intent evolves and the system adapts accordingly. This approach aligns with spec-driven development principles (Gupta, 2026) and recent AI-native engineering articles (OpenAI Developers, n.d.).### 6.5 Context and Constraint DependentAI-native delivery engines are context and constraint dependent, meaning their behavior changes based on the information they receive at runtime. The engine may use that behavior to deliver user-facing outcomes. Behavior emerges from the interaction between prompts, models, constraints, and contextual data supplied through retrieval pipelines (NB, 2026).Context and constraints shape behavioral intent. Stakeholder transcripts, domain rules, and operational constraints define the expected behavior. Retrieval pipelines inject domain knowledge. Structured and unstructured data are supplied to models at runtime to ensure accurate, grounded outputs. Testing validates behavior under different contexts. Integration ensures correct context flow. APIs, services, and orchestration layers deliver the right context to the right model at the right time.IntDEx assumes that AI-native engines operate mainly on context injection rather than model training. The latter is also endorsed assuming the capability exists and that organisational data privacy policies allow it.### 6.6 Model-AdaptiveModels evolve, drift, and differ in capability. AI-native delivery must therefore adapt continuously to model changes. IntDEx treats models as interchangeable components whose capabilities shape feasibility and whose behavior must be evaluated for fitness.The engine should optimize prompts for specific model families. Multiple models may be assessed to determine which produces the desired behavior. Model routing (the system that decides what model will be used) adapts to performance, cost, and latency. Drift detection identifies degradation in model behavior. When new model versions are released, prompts, constraints, and architectural components may need to be updated to maintain alignment. This adaptive stance reflects contemporary AI-native engineering patterns and ensures that digital solutions remain robust as underlying models evolve.Models evolve, drift, and differ in capability. AI-native development must adapt to model changes continuously (Debois, 2026).### 6.7 Human-Augmented and Evidential independenceIntDEx assumes that AI extends human skills and knowledge. IntDEx embeds human-in-the-loop oversight, ethical controls, and continuous monitoring for drift and hallucinations to ensure AI safety (Amanda McGrath, 2026), responsible and reliable AI-native delivery. Human experts remain key to decision-making, validating outputs and intervening when system behavior deviates from intent. Ethical safeguards constrain model actions through policy-aligned prompts, domain rules, and contextual safeguards. Continuous monitoring detects behavioral drift, hallucinations, and degradation across model versions, triggering corrective actions such as prompt refinement, constraint updates, or model change. Together, these mechanisms form an oversight layer that aligns with recent AI governance and responsible AI engineering practices and regulation like the EU AI Act (European Parliament, 2025).#### Human oversight as an enforceable control, not an intentionIn IntDEx, human augmentation is only considered implemented when it is expressed as a checkpoint, is recorded in a machine and human readable log, and can be verified. IntDEx defines human oversight through four concrete elements: mandatory checkpoints, accountability, a decision record, and a blocking rule.Mandatory human checkpoints. A human decision is required, and the engine must not proceed autonomously, in each of the following situations:1. Major Intent creation and intent change, because intent is the human-owned specification and an engine that authors its own acceptance criteria has no external reference.2. Any change that alters scientifically user-visible behavior, a business rule, a route, or a data contract.3. Any change affecting security, authentication, authorisation, privacy, or personal data handling.4. Any additive or destructive database schema change.5. Major release approval, based on the release checklist and validation evidence.6. Any breach of a cost or risk threshold defined in the engine manifest.7. Any case where the engine reports that it cannot validate its own output.Accountability. Each checkpoint is owned by a named role drawn from the IntDEx role model: the Delivery Manager owns scope, release, and cost or risk escalation decisions, the Architect owns security, privacy, and compatibility decisions, the Tester owns acceptance of behavior validation evidence, the Developer approves engine’s actions. The engine holds no approval authority for major checkpoints.Decision record. Every checkpoint produces an entry containing the reviewer identity, the role, the decision taken, the rationale, the artefacts inspected, the linked evidence, the change severity, and a timestamp. The record is written latest-first and is machine and human -readable so that it can be audited. The engine may create entry, but the engine must not populate the human-validation field. Only a human reviewer may set it.Blocking rule. Unresolved human checkpoints accumulate as a debt that may halt the engine. When the number of unvalidated high-severity checkpoints exceeds the threshold defined in the engine manifest, the engine must stop and report the blocking condition rather than continue.Human oversight is not that the human reviewer only reads a summary produced by the same engine that performed the work. IntDEx requires the reviewer to inspect the underlying artefacts and evidence, not the engine's self-description of them, and the decision record must name which artefacts were inspected.Non-mandatory checkpoint could also be implemented, these should not block the engine from progressing and can be performed asynchronous., In case of findings during these checkpoints the user should instruct/promt the engine to roll back.**Maturity** **relaxation**The HILT mandatory checkpoints could be relaxed based on the maturity of the AI delivery engine (models, practices, context understanding etc.)**The reviewer is also a failure mode** **and it may get** **worse** **as** the better the engine gets, because well-formatted, confident output is harder to challenge than poor output. A checkpoint register in which almost everything is approved is more likely to be evidence of rubber-stamping than of quality.IntDEx names four reviewer-side failure modes and asks teams to design against them rather than advice against them:- **Automation bias.** Accepting a machine recommendation without independent examination.- **Anchoring.** A reviewer shown the engine's conclusion first evaluates that conclusion instead of the question. This is the reason acceptance criteria must be confirmed before implementation exists, not only for traceability.- **Maintaining Skills.** A team that has stopped producing artefacts itself gradually loses the expertise needed to judge them.These are stated as a known limitation, not a solved problem. IntDEx does not claim that human checkpoints reliably catch machine errors. It claims that a checkpoint that is recorded, attributed, and inspectable can be audited, whereas informal review cannot. Teams should periodically sample approved checkpoints and re-examine them independently and treat a rejection rate of zero as a finding that needs further investigation.#### Independently VerifiedWhen the same model authors the intent, the prompt, the implementation, the test, and the test result, a misunderstanding may be reproduced consistently through all of them. Traceability then appears complete, coverage appears full, and validation appears good, while the entire chain is wrong. Traceability between artefacts written by one author may be just internal consistency, not verification, and internal consistency is exactly what a confidently mistaken model produces.This risk is specific to AI-native delivery and grows precisely as adoption succeeds. The more the engine is trusted to produce artefacts, the fewer independent references remain against which its output can be checked. Frameworks that rely on artefact completeness as the only indication for correctness may be unsafe.IntDEx addresses this through the suggestion of evidential independence: a claim of correctness is only admissible if its evidence originates from a source other than the model that produced the work being judged. In practice IntDEx suggests four admissible sources of independent evidence, in descending order of strength:1. Executed evidence. Output from a real execution, such as an execution exit code, a test result, a build log, an HTTP response, or a database state check. This is the strongest form because it is produced by the system rather than described about the system.2. Human evidence. Acceptance criteria validated by a human before implementation, or a human review decision recorded at a checkpoint.3. Cross-model evidence. Validation performed by a model other than the authoring model, ideally from a different provider family. This is advisory and reduces, but does not eliminate the risks,4. Prior-artefact evidence. Comparison against a baseline artefact version that pre-exist the current change.A statement by the authoring model that a check passed carries no evidential weight and must be recorded as unverified rather than as a pass. However, the team may decide, depending on the criticality of the evidence, the level or desired automation and the absence of alternatives to allow a same agent/model feedback loop. Not that until a gate's blocking path has been demonstrated, its passing output carries no assurance.The same suggestion goes for cost accounting. A cost figure estimated by the model about its own consumption is a self-assertion, while it might be a useful indicator, must be labelled as unverified self-report until it is checked against provider telemetry.## 7 IntDEx Delivery FlowThe IntDEx delivery flow describes AI-native delivery phases designed to support teams focusing on an AI-Native approaches. The delivery flow consists of phases where each phase is revisited as needed, reflecting the non-linear and emergent nature of AI-native development.### 7.1 Phase 1: Intent Definition**Purpose**Intent definition establishes the behavioral goals of the digital solution. Unlike traditional requirements analysis, which focuses on functional specifications, IntDEx defines requirements in terms of desired behavior, constraints, and acceptance criteria.**Process**Stakeholder documents, interviews, and domain knowledge are transformed using AI and other deterministic tools into intent, which expresses the expected behavior in domain-specific terms. The intent should be validated by Human. For critical intents (e.g. security, legal compliance, data privacy) humans can be authors and AI-validation may be used. Prompts then operationalise this intent into executable instructions for the AI delivery engine.**Output**The output of this phase is a set of initial intents that define the behavioral objectives of the digital solution. These specifications serve as the foundation for prompt development and generative creation.### IntDEx changes| **Closest existing practice** | **What IntDEx adds** | **What IntDEx changes** || --- | --- | --- || Requirements elicitation, user stories, Behavior-Driven Development scenarios, spec-driven development | Intent as a versioned artefact with its own baseline for regression detection, risk-aware constraints (cost, sovereignty, ethics) declared up front, not in a later Non-Functional Requirement test | Requirements are behavioral outcomes in domain language, not functional specs, intent is AI-authored by rule and validated by human with exceptions. Intent validation is a mandatory blocking HITL checkpoint. As a change for traditional agile practices IntDEx suggest that the work can start from intents. Epics, features, user stories while they can be kept if essential for the organisation/teams, they are not essential for IntDEx. |### Human vs. AI delivery engine in IntDEx| **Human does** | **Engine does** | **Who decides** || --- | --- | --- || Elicits stakeholder needs: Validates the intent, sets constraints (ethical, cost, sovereignty, domain rules); Validates acceptance criteria before implementation | Assists drafting and structuring; Drafts the Intent, operationalises intent into prompts. Validates human created artefacts. Writes acceptance criteria before implementation | Human**.** Intent validation is a blocking checkpoint. |### 7.2 Phase 2: Delivery Engine Architecture**Purpose**Delivery Engine Architecture defines the structural design of the AI-native delivery engine. It establishes how models, prompts, retrieval mechanisms, safeguards, and integration points work together to produce reliable behavior. The delivery engine may be owned by the product delivery team or provided as a shared organizational capability.**Process**The Architect designs, implements, and monitors the delivery engine architecture, ensuring that it is fit for purpose and compliant with security, data governance, sovereignty, ethics, interoperability requirements and cost constraints. The Delivery Manager evaluates and accepts the solution.**Output**The output of this phase is a fully integrated delivery engine capable of producing consistent behavior across the digital ecosystem. The architecture includes model routing that selects the appropriate model for each behavioral requirement, retrieval pipelines that supply contextual knowledge, and API integrations that ensure consistent behavior across endpoints. It incorporates observability mechanisms for detecting behavioral anomalies in real time, delivery engine-level prompts that define global rules and tone, task-level prompts that encode specific behaviors, safeguards that enforce ethical, safety, sovereignty and cost constraints, and context blocks that inject domain-specific information at a level below user intents and prompts.### IntDEx changes| **Closest existing practice** | **What IntDEx adds** | **What IntDEx changes** || --- | --- | --- || Solution architecture, Agile product management, MLOps/LLMOps platform design, reference architectures | Explicit separation of the AI delivery engine from the digital solution, each governed and versioned independently, four mandatory instruction artefacts (Manifest, Governance, Framework Reference, IDE/agent instruction set) as enforceable constraints. Sovereignty, ethics and cost as first-class architectural criteria alongside security. | Architecture is a recurring phase, not an up-front stage, the Architect owns model registries, routing, retrieval and rollback with the same discipline as application code |### Human vs. AI delivery engine in IntDEx| **Human does** | **Engine does** | **Who decides** || --- | --- | --- || Architect defines stack, routing policy, retrieval design, safeguards; assesses sovereignty, security, cost; DM accepts | Generates configuration, instructions, monitors and reports on itself | Human**.** Security, privacy, sovereignty, and compatibility decisions are Architect-owned |### 7.3 Phase 3: Generative Creation**Purpose**Generative Creation replaces traditional implementation. Instead of writing code, developers generate and refine behavior through prompts and models. This can include prototyping and proof of concepts.**Process**Developers and analysts engage in pairs to interact with the engine, collaboratively refining messages to the engine. Prompts can be produced by the AI engine following the messages to the engine and then be validated by humans. Prompts for critical intents can be authored by humans and validated by the AI engine. Multiple models are tested to identify the best behavioral fit. Constraints are tuned to eliminate undesired behavior, and prototypes are compared based on behavioral quality.Pair interaction is strongly recommended. If there is no analyst available, it could be a developer and a tester. However, if team plans and resources do not allow for pair interaction it can be applied either to more complex intents only or completely dropped. In the latter case more time may be required for validation and business user input.**Output**The output is a set of behavior prototypes that demonstrate how the AI delivery engine responds to prompts, constraints, and context, and how those responses shape the digital solution delivered to users.### IntDEx changes| **Closest existing practice** | **What IntDEx adds** | **What IntDEx changes** || --- | --- | --- || Implementation, pair programming, AI-assisted coding | Pair interaction (Developer + Analyst, or Developer + Tester) as the default working mode with AI validation. Multi-model comparison to select behavioral fit, prompt versioning rule, new intent means a new prompt file, same intent with better execution means a new version. Deterministic Data captured back from the model post-implementation to raise reproducibility | Coding is done by the AI delivery Engine. Humans concentrate on interaction authoring and refinement, the output is behavior prototypes, not only code, messages, intents and prompts are the key engineering asset |### Human vs. AI delivery engine in IntDEx| **Human does** | **Engine does** | **Who decides** || --- | --- | --- || Pair-interaction with the engine; steers, rejects, refines; compares candidate models on behavioral fit | Generates code, configuration, documentation, behavior prototypes. Records Deterministic Data | Shared**.** Human sets direction, engine produces; Developer may prepare but may not approve |### 7.4 Phase 4: Behavior Validation**Purpose**Behavior Validation ensures that development engine’s outputs are correct, safe, ethical, and aligned with stakeholder expectations. Testing focuses on behavior rather than code.**Process**Testers evaluate outputs for correctness, bias, hallucinations, constraint violations, drift.AI agents assist in automated evaluation. The engine should trigger self-assessments on major engine changes. while human-in-the-loop (HITL) testers validate domain-specific correctness (see AI Engine Monitoring). User Acceptance Testing (UAT) is the behavior validation from the end-user/sponsor. IntDEx recommends acceptance behavior validation in pairs e.g. User and Tester or User and Analyst. Intents, prompts and contracts can be redefined during paired acceptance behavior validation. UI/UX, Security, compliance tests are performed using automated tools before release. Human testers validate results.**Output**The output is a validated set of behaviors, along with identified issues and refined intents, prompts or constraints.### IntDEx changes| **Closest existing practice** | **What IntDEx adds** | **What IntDEx changes** || --- | --- | --- || Quality Assurance, UAT, test automation, evals (automated tests that measure an AI system’s behavior) | Continuous Behavior Validation, evaluations fire on any engine configuration change (prompt, constraint, context, model swap), gating on pass rate. Pair Behavior Validation (End-user + Tester / Analyst / DM). Evidential independence with a ranked hierarchy, executed > human > cross-model > prior-artefact. Non-functional requirements testing is automated (security, UI/UX) in the testing phase. | Testing targets behavior, not code. Correctness, bias, hallucination, constraint violation, drift. A self-reported pass by the authoring model is inadmissible and must be recorded as UNVERIFIED. |### Human vs. AI delivery engine in IntDEx| **Human does** | **Engine does** | **Who decides** || --- | --- | --- || Testers accept or reject evidence. End-users do pair acceptance validation. Judges bias, hallucination, domain correctness | Runs automated tests and engine evaluations. Reports failures, drift, constraint violations, output anomalies | Human**.** The engine's own report is not evidence. A self-assessment pass is UNVERIFIED |### 7.5 Phase 5: Release**Purpose**Packaging bundles of all behavioral artefacts required for deployment Release package contain behavior bundles and delivers a new version of the digital solution.**Process**Release packages include prompt bundles, model versions, constraint versions, validation metadata, deployment notes, context versions to ensure traceability and reproducibility. Human release approval is mandatory for production. This is also in line with the EU AI Act.IntDEx suggests:-Progressive functionality release. New or changed behavior reaches a limited user base first, with the exposure widened only after the observed behavior matches expectations. Release in small chunks. Release flags can be implemented easier with the AI delivery engine.- Rollback in IntDEx means returning to a prior behavioral bundle, the previous prompt, constraint, context and model versions together, not only redeploying a previous code build.- A tested rollback path. Reverting must have been performed at least once in a test environment. An untested rollback procedure is an assumption.**Output**The output is a complete release package ready for deployment, containing all artefacts required to reproduce system behavior.### IntDEx changes| **Closest existing practice** | **What IntDEx adds** | **What IntDEx changes** || --- | --- | --- || Release management, Software Bill of Materials (SBOM), DevOps packaging | Release bundles contain prompt bundles, model versions, constraint versions, context versions and validation metadata. Everything needed to reproduce behavior, not just to redeploy code. explicitly aligned to EU AI Act traceability | Release takes place after a readiness signal, not at an iteration boundary. The overall framework validation test is rebuildability from artefacts. HITL checkpoint owned by the DM |### Human vs. AI delivery engine in IntDEx| **Human does** | **Engine does** | **Who decides** || --- | --- | --- || DM approves the release against checklist and evidence | Assembles the bundle: prompts, model versions, constraints, contexts, validation metadata | Human**.** Releases are never self-deployed |### 7.6 Phase 6: AI delivery engine continuous Monitoring**Purpose**Continuous Monitoring tracks AI delivery engine behavior in production, identifying drift, degradation, and emerging risks (security, cost etc.).**Process**Monitoring includes drift detection, hallucination monitoring, constraint degradation tracking, data quality alerts, real-world usage analysis, AI safety. Continues automated validation should be used. AI-engine instructions with known deliverables can be passed on to the new version of AI-engine, the new results should be compared to previously accepted results. Phase 4 (behavior Validation) can trigger Monitoring ActionsThese signals trigger new delivery cycles, reflecting the non-linear nature of AI-native systems.**Output**The output is a set of monitoring insights that inform prompt updates, model adjustments, and architectural refinements.### IntDEx changes| **Closest existing practice** | **What IntDEx adds** | **What IntDEx changes** || --- | --- | --- || Application Performance Monitoring /observability; MLOps drift monitoring, Site Reliability Engineering | Monitoring extended beyond correctness to cost, sovereignty, ethics and AI-safety drift. Constraint degradation tracking (safeguards weakening over time). Breach of a cost or risk threshold is a mandatory blocking HITL checkpoint | Monitoring signals directly triggers new delivery cycles into any phase, it is an input to delivery, not a post-release operational concern |### Human vs. AI delivery engine in IntDEx| **Human does** | **Engine does** | **Who decides** || --- | --- | --- || Triages alerts. Decides what is a real deviation; owns cost and risk thresholds | Detects drift, degradation, cost and sovereignty breaches. Raises signals | Shared detection, human judgement**.** A threshold breach forces a checkpoint |### 7.7 Phase 7: Behavior Evolution**Purpose**Behavior evolution replaces traditional maintenance. Maintenance becomes the continuous evolution of both the AI delivery engine’s behavior and the user-facing behavior of the digital solution.**Process**Behavior evolves through:- new and updates intents, Intent versioning- development engine improvements, model upgrades, Model versioning and or AI delivery Engine Versioning- prompt refinement- constraint tightening- architectural adjustmentsThe AI delivery engine can be extended to monitor digital solutions operations. Findings create intents for human validation. Support requests can be triaged by the AI delivery engine and create intents and messages for the users. HITL oversight ensures ethical and domain-specific correctness throughout evolution.**Output**The output is an improved behavioral profile that reflects new data, new models, and new stakeholder needs.### IntDEx changes| **Closest existing practice** | **What IntDEx adds** | **What IntDEx changes** || --- | --- | --- || Maintenance, product backlog, continuous improvement | Versioned baselines across intent, prompt, constraint, model and engine so behavioral regression can be detected against a prior baseline. AI delivery engine (or an extension of the engine) is monitoring solution operation and receives support request | Maintenance is redefined as evolution of behavior, and evolution of the engine is tracked separately from evolution of the solution |### Human vs. AI delivery engine in IntDEx| **Human does** | **Engine does** | **Who decides** || --- | --- | --- || Decides what changes and why, re-validates intent | Updates Intents, applies changes, versions artefacts. Maintains rebuild evidence. | Human for the intent, engine for the implementation |### Cross-cutting changes (not phase-specific)| **Area** | **Existing practice** | **IntDEx position** || --- | --- | --- || Flow | Sprints / time-boxed iterations | Flow-based, readiness-triggered replenishment (Kanban-like); replenishment as needed, delivery planning monthly, status every 1–3 weeks, strategy quarterly. Sprint governance may coexist but is not essential. || Management | Separate Project and Product Management | Product Delivery Management as one unified practice || Team shape | Fewer, larger, technology-aligned teams | More, smaller (3–5), business-domain-aligned teams. Sized by business/regulatory complexity, not technical complexity || Human oversight | Review, sign-off, RACI | Defined blocking checkpoints, role-owned, with a decision record the engine may create but must not sign. Unvalidated high-severity checkpoints accumulate as debt that halts the engine || Evidence | Traceability matrices, coverage | Traceability and coverage are explicitly **not** substitutes for correctness evidence || Cost | Budget tracking | Per-delivery cost estimation and actuals visible to the whole team, owned by the DM. Model self-reported cost is an unverified self-report until checked against provider telemetry || Adoption | Full framework or nothing | Adopt what you need. A stated 5-item minimum threshold to claim IntDEx conformance, plus a throughput-maximisation mode (2–4 parallel AI streams, file-partitioned) for low-criticality work |## 8 Roles and ResponsibilitiesIntDEx is not reinventing SDLC roles however it proposes AI-native related responsibilities. AI-native delivery requires adjustments to traditional software engineering roles. IntDEx suggest a role model tailored to AI-native development, ensuring that governance, ethical oversight, and behavioral correctness are embedded throughout the delivery flow. Each role contributes uniquely to intent-centric design, behavior-driven engineering, and model-adaptive workflows. The “AI” or “IntDEx” label is assumed to every role e.g. IntDEx/AI Delivery Manger, IntDEx/AI Developer.Organisations may maintain their current role names and adjust the responsibilities as needed.Two structural rules bind this role model, and an IntDEx implementation that states roles withoutthem has described a courtesy rather than a control:1. **Every role is an approval authority or it is not a checkpoint role.** A role named here thatowns no blocking checkpoint, no artefact and no prohibition is decoration. Each role belowtherefore closes with what it may *never* do.2. **Authority is exclusive.** Exactly one role owns each checkpoint. Where two roles appear to ownthe same decision, the split is stated explicitly rather than left to negotiation, because twodocumented owners of one decision reliably produce nought.### 8.1 Business Owner (BO)**Purpose**The Business Owner defines and safeguards the business intent behind every AI-native developmentinitiative. Their core purpose is to explain the *why* behind a feature, product or change: thevalue it must deliver, the problem it must solve, and the constraints it must respect. The BusinessOwner ensures that every agent-generated artefact aligns with the organisation's goals, customerneeds, regulatory boundaries and long-term strategy. In an engine where implementation is cheap andnear-instant, intent is the only scarce input, and the Business Owner is its source.**Responsibilities**The BO defines the intent, constraints and success criteria that guide autonomous AI agents,clarifying the problem to solve, the target users, the expected outcomes, and the boundaries withinwhich the engine must operate: regulatory requirements, budget limits, risk tolerances andstrategic priorities. The BO authors or confirms acceptance criteria *before* implementation, whichis what makes the evidential-independence rule enforceable rather than aspirational. They ownprioritisation and sequencing, and they validate whether delivered functionality meets the businessintent. The BO gives the **business go/no-go for a release**; the Delivery Manager records therelease event. Together with the Risk Manager, the BO is a joint reviewer of any EC-07determination, because a judgement about unlawful or seriously harmful purpose must not rest with asingle individual.**May never**: author acceptance criteria and accept the resulting implementation in the same act;record a release event; waive an ethical constraint.### 8.2 Delivery Manager (DM)**Purpose**The equivalent of the Project/Product Manager role, ensuring that AI-native delivery activitiesalign with strategic objectives, governance requirements, initiative constraints (time, resources,cost) and stakeholder expectations, while monitoring AI risks and mitigation measures. This reflectsPMI’s emphasis on strategic value, governance, and stakeholder engagement (PMI, 2026).**Responsibilities**The DM uses AI-augmented tools to create project initiation documents, reports and monitoringartefacts. The DM validates AI-generated documents and artefacts through human-in-the-loopoversight, ensuring ethical and contextual correctness. The DM confirms the AI tools, agents andmodel usage strategy, ensuring alignment with organisational constraints and compliancerequirements. Reporting and project planning are performed using structured intents and prompts,reflecting the intent-centric nature of IntDEx. The DM owns **scope** and **cost/risk escalationwhere no dedicated Risk Manager exists**, and is the **named actor who submits the release event**to the lifecycle ledger once the Business Owner has given the business go/no-go. Because a releaseis a physical act the engine cannot observe, an unattributed release is indistinguishable from aself-declared one and is refused.**May never**: submit a release without a recorded Business Owner go/no-go; set `validated_by_human`on another role's checkpoint; accept a result the Tester has classified as `UNVERIFIED`.### 8.3 Architect**Role Purpose**The Architect designs the technical foundation of the AI-native development engine, ensuring thatprompt architecture, model routing, retrieval pipelines and integration points support the desiredbehavior.**Responsibilities**The Architect defines the technology stack, integration points and orchestration layer. They reviewprompt templates for architectural alignment and ensure compliance with security, data governanceand interoperability standards. Model management is a core responsibility: models are versioned,tested, deployed, monitored and rolled back with the same discipline as application code, and modelregistries, experiment tracking and automated evaluation pipelines are integrated into the deliveryflow. Third-party but also integral AI services are assessed for sovereignty, cost and security. TheArchitect owns **sovereignty, security, privacy and compatibility decisions**, and is the namedhuman who **adopts the engine's capability boundary and incident-autonomy artefacts** - the twoartefacts that define what the engine may do without a human, and which remain inert until a namedhuman signs them.**May never**: widen engine autonomy without a Major checkpoint; approve their own architecturalchange as validated; treat a sovereignty assessment as a monitoring control.### 8.4 Analyst**Purpose**The Analyst translates stakeholder needs into behavioral intent, replacing traditional user storieswith intents and prompts, and validates multiple delivery-flow artefacts.**Responsibilities**The Analyst defines epics, features and intents in terms of behavioral outcomes, and maintains workitems in tools that support intent and prompt storage and versioning. Requirements evolve throughgenerative creation, and Analysts collaborate with developers and testers in pair-interactionsessions to deliver the digital solution. The Analyst owns **intent creation and intent change**,which is a mandatory blocking checkpoint: an intent altered without a recorded human decisionsilently redefines what "correct" means for every downstream artefact and test. The Analyst alsoowns the `Work Item` join key, without which lifecycle and lead time cannot be computed.**May never**: approve an intent they authored in the same act; accept behavior validation evidence;reclassify a business constraint as a preference.### 8.5 Developer**Purpose**The Developer generates and refines system behavior through interaction with the delivery engine,prompts, models and constraints. This shifts the focus from traditional code-centric tobehavior-centric engineering.**Responsibilities**Developers validate and edit prompts, execute prompts to evaluate outputs, and engage in pairinteraction with analysts or testers. They ensure the link between intents and resulting behavior isclear through artefact versioning and interaction logs, and they maintain behavioral consistencyacross model updates. Developers test multiple models to identify the best behavioral fit,reflecting model-adaptive engineering practices, and advise the Architect on possible changes to theAI delivery engine.**May never**: approve any checkpoint. The Developer prepares evidence and proposals; approval restswith the owning role. This is the single most frequently violated rule in practice, because theDeveloper is usually the person closest to the change and the quickest to sign it.### 8.6 Tester**Purpose**The Tester validates solution behavior, ensuring correctness, safety and compliance. Testing focuseson behavior rather than code. The Tester validates the AI delivery engine itself as well as theproduct it produces.**Responsibilities**Testers use AI agents to evaluate outputs for correctness, bias, hallucinations and constraintviolations. They perform behavior validation with human-in-the-loop oversight, ensuringdomain-specific correctness. Testers identify drift, degradation and emergent behaviors that requireprompt or model updates, and work with end users and sponsors in pairs for acceptance behaviorvalidation. The Tester owns the **`verification_source` classification** of every recorded resultand is the **only role that may accept a result weaker than `executed`**. Accepting an `UNVERIFIED`result is a decision with a named owner, never a default that occurs because nobody objected.**May never**: round an `UNVERIFIED` result up to a `PASS`; accept the engine's own report asindependent evidence; validate behavior against criteria authored in the same generation act as theimplementation.### 8.7 Risk Manager**Purpose**The Risk Manager ensures that the AI-native delivery engine and the resulting digital solutionremain safe, ethical and compliant throughout their lifecycle. This role reflects PMI’s emphasis onrisk, ethics and governance. The tasks may be performed by the Delivery Manager where there is nodedicated Risk Manager.**Responsibilities**The Risk Manager monitors model drift, bias, hallucinations, cost, security, data leakage andsovereignty. They maintain risk logs, escalation paths and compliance documentation. Ethicalsafeguards are reviewed continuously, and new risks trigger updates to constraints, prompts androuting strategies. The Risk Manager ensures that governance checkpoints adapt to emerging riskpatterns (PMI, Cognilytica, 2025). Critically, the Risk Manager is the **named human who authors andadjudicates the ethical constraints EC-01 to EC-08**, and who records every EC determination -including a cleared false positive - as a Major checkpoint. Until that named review occurs, theconstraints whose detection is not mechanically decidable remain `UNVERIFIED`, and a passing ethicsgate confers no clearance whatsoever. **EC-07 is unwaivable and is reviewed jointly with theBusiness Owner.** The Risk Manager also owns **cost and risk thresholds and escalation**, and**sovereignty monitoring** (the Architect owns the sovereignty decision; the Risk Manager ownswhether it still holds).**May never**: waive EC-07; adjudicate EC-07 alone; accept an engine self-assessment as an ethicsreview; allow an ethics finding to be closed without a recorded determination.### 8.8 Optional specialisation: Data and Sovereignty OwnerWhere sovereignty, data residency or regulatory classification is central to the product rather thanincidental to it, organisations may split a dedicated Data and Sovereignty Owner out of theArchitect and Risk Manager roles. This is a specialisation, not an eighth authority: the decisionrights remain those described above, and the split must be recorded so that exactly one role stillowns each checkpoint.### 8.9 RACI matrixR = performs the work. A = accountable, single owner of the decision. C = consulted before thedecision. I = informed after it. The engine holds **no approval authority at any checkpoint** andappears in no cell; it prepares, records and blocks.| Decision / checkpoint | BO | DM | Architect | Analyst | Developer | Tester | Risk Mgr ||---|---|---|---|---|---|---|---|| Business intent, value, success criteria | **A/R** | C | I | C | I | I | C || Acceptance criteria authored before implementation | **A** | C | I | R | I | C | I || Prioritisation and sequencing | **A** | R | C | C | I | I | C || Intent creation or intent change | C | C | C | **A/R** | C | C | I || Epic / feature / prompt authoring | I | I | C | **A/R** | R | C | I || Technology stack, integration, orchestration | C | C | **A/R** | I | C | I | C || Model selection, routing, versioning, rollback | I | C | **A** | I | R | C | C || Sovereignty decision (residency, provider, service) | C | C | **A/R** | I | I | I | C || Sovereignty monitoring (does it still hold) | I | C | C | I | I | C | **A/R** || Security, authentication, authorisation change | I | C | **A** | I | R | C | C || Privacy or personal-data change | C | C | **A** | I | R | C | C || Database schema or data-contract change | I | C | **A** | C | R | C | I || User-visible behaviour, route, business-rule change | C | **A** | C | R | R | C | I || Implementation and artefact updates | I | I | C | C | **A/R** | I | I || Behaviour validation evidence acceptance | C | C | I | C | R | **A/R** | I || `verification_source` classification | I | I | I | I | C | **A/R** | C || Accepting a result weaker than `executed` | I | C | C | I | I | **A** | C || Cost and risk threshold definition | C | C | C | I | I | I | **A/R** || Cost or risk threshold breach escalation | I | **A** (if no RM) | C | I | I | I | **A/R** || Risk log ownership and closure | I | C | C | I | I | C | **A/R** || EC-01 to EC-08 authoring and adjudication | C | C | C | I | I | C | **A/R** || EC-07 determination (unwaivable) | **A** (joint) | I | C | I | I | I | **A/R** (joint) || Engine capability boundary adoption | I | C | **A/R** | I | C | C | C || Incident-autonomy thresholds and plan pre-approval | I | C | **A/R** | I | C | C | C || Engine constraint / gate change | I | C | **A** | I | R | C | C || Release business go/no-go | **A/R** | C | C | I | I | C | C || Release event submission (named actor) | C | **A/R** | I | I | I | I | I || Release checklist completion | I | **A** | R | R | R | R | R |Two cells are deliberately joint rather than single-owner: EC-07, because a judgement aboutunlawful or seriously harmful purpose is too consequential to rest on one person; and cost/riskescalation, where the Delivery Manager absorbs the Risk Manager's accountability only when no suchrole exists. Every other decision has exactly one A.## 9 Expanding and clarifyingThis chapter further clarifies, emphasises and expands some of the concepts already mentioned earliein the paper.### 9.1 Product Delivery Management**Flow-Based Delivery**Traditional Agile methodologies rely on boxed iterations such as sprints. However, AI-native development can support even faster flow-based delivery, where work progresses continuously based on readiness signals rather than fixed iteration boundaries (Denis Dennehy, 2018). IntDEx adopts a Kanban-inspired approach, emphasizing continuous flow, work-in-progress limits, and readiness-based replenishment, deployment and release.Replenishment meetings determine which work enters the AI delivery engine next, based on clear policies, capacity, and risk. Delivery planning occurs monthly, confirming readiness signals for release. Delivery status updates occur every one to three weeks, focusing on lead time, risks, capacity, and strategic alignment. Strategic planning occurs quarterly or annually, aligning medium-term roadmaps with organizational objectives.This flow-based approach aligns with the non-linear nature of AI-native development, where behavior evolves continuously and cannot be constrained by rigid iteration boundaries.**Pair** **Interaction,** **Pair Behavior Validation**Pair interaction and Pair Behavior Validation are considered important IntDEx practices where two roles collaborate to refine interaction (messages to the AI Delivery Genine) or to validate (test) the resulting behavior. Common pairings include Developer + Analyst, Developer + Tester, and Delivery Manager + Architect. Pair interaction improves accuracy, reduces hallucinations, accelerates refinement, and enhances shared understanding (ByteByteGo, 2026). However, an IntDEx team may decide to reduce pair interaction at times as needed in order to ensure smooth operation. Pair Behavior Validation (End-user + Tester, End-user + Analyst, End-user + DM) improves end-user/sponsor experience, it captures issues beyond functionary problems (e.g. indications of delivery engine disfunction) and can facilitate simultaneous intent and prompt refining.**Prompt Versioning and Bundling**Prompts are versioned (e.g., v{v}, v{v}.1, v2) and stored in repositories such as Git or work item fields. Prompt bundles group related prompts by feature, epic, release, or integration point. This ensures reproducibility, traceability, and governance.A simple rule to identify when a prompt should be stored as a new prompt or as a new version on an existing prompt is-A prompt becomes a new prompt when the intent changes. Intent change means new prompt file-A prompt becomes a new version when the intent stays the same, but the execution improves Intent stays the same means the same file is versioned.E.g.Prompt A, v{v}: “Generate architectural options.”Prompt A, v2: “Generate architectural options using constraints X, Y, Z.”Prompt B v{v}: “Refactor code”Prompt C v{v}: “Generate test cases”**Behavior Testing**Behavior testing tests the results of the prompt and evaluates determinism, constraint adherence, hallucination, bias etc. Behavior should be tested after each prompt, individually and then again as a whole before a release or after the development of a major functionality.**Deterministic Data**Because intents and prompts when created in most cases will not include all the implementation details and in order to increase the determinism of the AI delivery engine should record some Deterministic Data provided by the AI model/service after the implementation of the prompt. This additional information will improve the reproducibility of the solution. This data should be machine and user readable. It shouldn’t be code.### 9.2 Risk Controls**Nature of AI-Native Risks**AI-native delivery introduces evolving risks such as model drift, hallucinations, bias, data leakage, automation failure, cost, security and sovereignty. These risks arise because AI systems generate behavior, not deterministic logic, and that behavior changes as intents, prompts, models, and data evolve. In addition, most organisations will have to use third party AI endpoints, and this may increase cost, security and sovereignty risk.**IntDEx Risk Controls**IntDEx recommends preventive, detective, and corrective controls directly into each delivery flow phase. This ensures risks are not only detected but structurally prevented. Constraint-first prompt design: Prompts embed or reference domain rules, safety constraints, and formatting requirements to prevent drift and hallucinations before execution. Cost, security and sovereignty should be documented and monitored as part of the AI delivery engine monitoring. Use of deterministic tools and libraries for code quality checks.**Untrusted content and instruction injection**An AI delivery engine that navigates multiple folders, reads repository content, retrieves documents and applies patches will process text that no member of the team wrote. Any such text may contain instructions aimed at the model. Source code comments, dependency README files, issue text, web pages, retrieved documents, and even prior engine output are all untrusted input in this sense.Constraint-first prompt design does not mitigate this. Constraints placed in the prompt and instructions injected through retrieved content arrive at the model as the same kind of token, so a rule stating "ignore malicious instructions" is itself only another instruction competing for attention. IntDEx therefore treats instruction injection as an architectural problem, not a prompting problem, and requires the following:- **Content is data, never instruction.** Retrieved and file-derived content is delimited and labelled as untrusted on entry. An instruction discovered inside untrusted content is reported as a finding, never executed.- **Capability limitation over** **influencing** **the** **engine.** The engine's ability to do harm is bounded by what its credentials, tools permit and deterministic scripts, not by what it has been asked not to do. Least privilege applies to file system scope, network, package installation, credential access and command execution. Where the engine can execute commands, it does so in an isolated environment.- **Reduce the maxim possible damage** **before trust.** Before granting a new capability, the team records what the worst outcome would be if the engine were fully under an attacker's control while holding it. If that outcome is unacceptable, the capability should not granted.- **Export** **is a control point.** Export of repository content or secrets requires an outbound path. Outbound destinations available to the engine are specific and restricted.- **Injection** **(intent, prompt etc.)** **is a standing risk-log entry.** with a named owner, not an incident category discovered after the fact.This risk grows with autonomy. It is highest exactly where IntDEx is most useful, in multi-file, multi-folder agent operation, so it is treated here as a first-class delivery risk alongside drift and hallucination.**Security of generated code and dependencies**IntDEx stats that behavior is the primary engineering artefact, and that users no longer focus on code. It does not follow that nobody inspects code. A vulnerability is a behavior, it is simply a behavior that no intent requested, and no acceptance criterion describes, which is precisely why behavior validation derived from intent will not find it.Behavior validation answers "does it do what was asked". It does not answer, "what else does it do". IntDEx therefore requires a separate inspection of what the engine produced:- **Static analysis** of generated code, run automatically, with findings treated as blocking until checked.- **Dependency and supply-chain scanning.** Generated code introduces dependencies that no human chose. Known-vulnerability checks run on every dependency change, and unknown or unmaintained packages are escalated rather than accepted.- **Secret scanning** over both the repository and the engine's own artefacts, including prompts, contexts and logs.- **Default-credential and insecure-default detection.** A seeded credential, a permissive default, or a disabled security control must fail a check. Reporting it in prose in a completion summary is not a control. A human reading a long summary will not reliably act on it.- **Third-party and** **licence** **review** for anything the engine introduces into the product.These checks are deterministic and run without a model in the decision path, which makes their output ana acceptable executed evidence. Where a check cannot be automated, the gap is recorded rather than assumed closed. For regulated projects/domains such as public health, finance and government, this inspection might not optional.**Preventive Controls** (How risks are avoided)Intent DefinitionBehavioral clarity: Well-defined behavioral intents reduce ambiguity and drift.Risk-aware constraints: Sensitive data, ethical rules, and domain boundaries, cost constraints, sovereignty assessments etc. are defined upfront.Prompt ArchitectureConstraint-driven prompt design, Prompts embed safety, formatting, and domain rules to prevent hallucinations and bias.Structural templates: Consistent prompt patterns reduce the risk of unpredictable results and automation failure.Example-based grounding: Curated examples reduce hallucination risk.Generative CreationControlled generation workflows: Multi-step creation pipelines reduce variability and enforce constraints.Safe model routing: Tasks are routed to models with appropriate safety and reasoning profiles.Behavior ValidationValidation gates: Outputs are checked against domain rules, constraints, and expected formats before release.Hallucination filters: Outputs are compared with known data sources to prevent fabricated content.ReleaseGuarded deployment: Only validated behaviors are released, preventing failures in production.Data minimisation: Sensitive information is stripped from prompts and outputs.Continuous MonitoringBehavior drift detection: Monitors changes in output patterns across time and model versions.Bias monitoring: Checks for demographic, cultural, or domain bias in real-world usage.Leakage scanning: Automated checks ensure no sensitive data appears in outputs.Automation failure detection: Monitoring the reliability of AI-driven processes.Cose, security, Sovereignty assessments: Monitor AI delivery engine changes that may affect compliance with the predefined requirements.Intent EvolutionControlled updates: Changes in business rules or domain context trigger structured intent updates, preventing drift.Versioned behavioral baselines: Each intent version defines expected behavior, enabling regression detection.**Corrective Controls (How risks are resolved)**Prompt refinement: constraints, examples, and structure are updated when behavior deviates.Generative pipeline adjustments: Routing, steps, or models are changed when failures occur.Constraint reinforcement: Additional safeguards added when degradation is detected.Intent updates: Behavioral definitions evolve when domain rules or business context change.**Governance & Escalation**Risk work items and escalation pathways ensure issues are documented, triaged, and resolved. Ethical safeguards evolve as new behaviors emerge, aligned with organisation’s AI governance practices.### Incident Autonomy and ResponseIntDEx constrain autonomy on two criteria, and an action is permitted only when both criteria permit it.**Criteria A, the severity, decides when.** Detection is performed by a deterministic script over a metric with a stable rolling baseline, with no model involved in the decision. Deterministic detection is the only point in the delivery flow where evidential independence is achieved structurally rather than by requesting human attention. At high severity the engine records only. At medium severity the engine may read, analyse and draft an intent, but its diagnosis is a hypothesis carrying model-self-report-unverified, never a confirmed root cause. At low severity the engine may act, but only by opening a change for review or by executing a plan that a human approved in advance.**Criteria B, the change** **types, decides what.** A statistical threshold measures deviation. The change types are the mandatory human checkpoints already defined in this framework. Where a change affects security, authentication, authorisation, privacy, a database schema, a data contract, user-visible behavior, a release, or the engine's own constraints, the engine may prepare a change for review but may never execute it autonomously, whatever the severity value.**Pre-approval is a stronger form of oversight than real-time approval.** A plan authorised by a human who is not under incident pressure reflects better judgement than an approval extracted from an ad-hoc under stress reviewer. IntDEx therefore treats planned pre-approval as a major checkpoint, recorded ahead of the incident, and treats any subsequent change to a plant as canceling that approval.**Autonomy scope is itself a governed artefact.** Severity thresholds and the plan list determine how much the engine may do without a human.Every action, finding and triage decision is logged with a timestamp. A human checks every finding above the logging severity. When a fix is implemented, an evaluation for that incident type is added to the evacuation suite, so the same failure cannot recur silently. This is a maturity relaxation of the checkpoint model.### Cost controlThe use of AI services will intrude addition costs to digital solution delivery initiatives. It is important that the product delivery team has strong cost estimation and control mechanisms. IntDEx suggests that cost information should be available to every team member of the product delivery team. The AI delivery engine should provide a mechanism for cost estimation and actual costs as the engine is used to deliver the product. Responsible for the cost control is the DM. The architect and the developer may provide advice on cost reduction mechanisms and AI delivery engine optimisations.### Governance**Governance Principles**Governance ensures accountability, auditability, compliance, and delivery flow controls (PMI, 2026) . AI-native delivery requires governance mechanisms that operate continuously rather than at predefined checkpoints.**IntDEx Governance Mechanisms**IntDEx incorporates governance through:- Intent versioning- Prompt versioning- Ai delivery engine / Model versioning- Audit logs (e.g. chat messages)- Compliance checks- Delivery flow controls- Governance context files- Constraints- Delivery engine configuration- Use of deterministic tools in the delivery flow- HILT safeguardsIntDEx delivery engine should implement artefacts hierarchy and escalation methods and describe HITL checkpoints. These mechanisms ensure that all AI-generated artefacts are traceable, reviewable, and governed throughout the delivery fflow.### 9.4 Ethics**Ethical Principles**IntDEx adopts, after contextual adjustment, the principles outlined in the EU’s Ethics guidelines for trustworthy AI (High-Level Expert Group on AI presented Ethics Guidelines for Trustworthy Artificial Intelligence, 2019). IntDEx requires that all components of the AI Delivery Engine and the resulting Digital Solution are:**lawful:** Compliance with applicable laws and regulations. E.g. the EU AI Act (European Parliament, 2025).**Ethical:** Respecting fundamental ethical principles, human values, and societal norms.**Robust:** Technically reliable and resilientIntDEx implements the requirements of the EU’s Ethics guidelines for trustworthy AI via its core characteristics and delivery flow phases as follows:- Human agency and oversight, ensured through the IntDEx Human-Augmented characteristic, which embeds human-in-the-loop review and decision authority.- Technical robustness and safety, supported by the IntDEx Delivery Engine Architecture and the Continuous Monitoring phase, which detect and mitigate failures, drift, and unsafe behavior.- Privacy and data governance, enforced through the Delivery Engine Architecture and Continuous Monitoring, ensuring responsible data handling and preventing leakage.- Transparency, achieved through versioning of Intent, Prompt Architecture, Model routing, and the Delivery Engine Architecture, enabling traceability and auditability.- Diversity, non-discrimination, and fairness, embedded in the Delivery Engine Architecture and monitored continuously to detect and correct biased behavior.- Societal and environmental well-being, supported by the Delivery Engine Architecture and the Human-Augmented characteristic.- Accountability, defined through clear IntDEx roles and responsibilities and reinforced by the Continuous Monitoring phase, which provides evidence for oversight and governance.IntDEx expands the above with- AI Safety, Sovereignty and Cost transparency, supported by the Human-Augmented characteristic, IntDEx Delivery Engine Architecture and the Continuous Monitoring phase that should ensure the AI delivery engine, can be controlled and act predictably.**IntDEx Ethical Safeguards**IntDEx embeds ethical safeguards directly into prompt architecture, model routing, and behavior validation. Constraints enforce ethical boundaries, preventing harmful outputs. HITL oversight ensures that ethical considerations are reviewed continuously. Ethical compliance is validated during behavior testing and monitored in production.### 9.5 Human-in-the-Loop (HITL) Oversight**Role of HITL**Human-in-the-loop oversight is essential because AI systems cannot reliably detect all forms of bias, hallucination, or contextual error (PMI, 2026). Humans provide domain expertise, ethical judgment, and contextual understanding that AI lacks. Moreover, the use of AI services requires security, sovereignty and cost control. IntDEx enforces the use of AI also for the latter set of assessments and monitoring, however, at the top of every control pyramid there should be a human.**IntDEx HITL Implementation**HITL checkpoints exist in every delivery flow phase. HITL oversight ensures that system behavior remains aligned with stakeholder expectations, ethical standards, and domain requirements. IntDEx roles and responsibilities further clarify human involvement and responsibility.### Automated Continuous Behavior Validation (ACBV)IntDEx requires that behavioral artefacts are validated continuously. Change to the Delivery engine such as configuration, prompts, constraints, contexts, and model selections are versioned artefacts that evolve, each change must be treated as having a behavioral impact. ACBV aims to extend the Behavior Validation phase into an automated, continuous-driven mechanism that executes non-interactive behavioral tests whenever the delivery engine configuration changes (Claude Academy, 2026).ACBV operate as the AI-native equivalent of stage gates. When a manifest file is rewritten, a constraint updated, or a model swapped, the evaluation suite verifies whether the delivery engine still produces behavior aligned with intent. Each evaluation is derived from real tasks, incidents, or domain rules and becomes a permanent regression test. Configuration changes are gated on evaluation pass rates, ensuring that behavioral quality cannot degrade silently.ACBV complement human oversight by providing additional mechanisms for automated evals to detect drift, hallucinations, constraint gradual weakening, and unintended behavioral shifts before they reach users. This mechanism operationalises IntDEx characteristics of model adaptivity, behavior-driven engineering, and continuous intelligence loops, ensuring that the delivery engine remains stable, governable, and reproducible across its entire flow.**Things to keep in mind:**Behavioral evaluations of a non-deterministic system vary between runs even when nothing has changed. IntDEx therefore recommends that evaluation gating distinguishes a real regression from run-to-run variation before blocking. Establish the expected variation for a suite/gate on a change that exceeds it rather than on any change.### Behavior-quality vs Throughput-maximisationIntDEx promotes a behavior-quality maximisation model with suggestions such as pair-prompting and HITL checkpoints. However, for low criticality initiatives, proof of concepts or where the goal is fast delivery even by increasing quality risks, a throughput maximisation model can be used. In this case any human resource can work in parallel AI streams (2-4) to reduce waiting times. It is important to implement a delivery plan that groups the tasks depending on the files/artefacts they edit. A stream cannot edit the same files/artefacts that another stream is also editing.## IntDEx tailoring and adjustments for adoptionOrganisations and teams may adopt IntDEx practices at different levels depending on their AI-maturity, operational capacity, trust in AI systems, and risk tolerance. IntDEx does not require full-scale transformation from day one, instead, it defines a minimum adoption threshold that ensures the delivery approach is genuinely AI-native.At a minimum, teams must:1. Use AI service messaging/prompts as the primary interface for work rather than traditional requirements writing, coding, documentation, reporting, design etc. (Generative Creation replaces traditional implementation).2. Create at least intent, prompt, constraint, context test prompt and test result artefacts. These should be versioned, and the solution is rebuildable from them.3. Maintain at least one mandatory blocking human checkpoint for delivery-engine performance confirmation before use and major user-facing deliverables (Human oversight is an enforceable control, not an intention).4. At least one item of admissible evidence (executed or human-validated result) before delivery to production.5. Deterministic or different family AI agent plus some deterministic methods or human Common Vulnerabilities and Exposures (CVE) check before delivery to production.6. The engine's capability boundary is written down. What the engine can read, execute, install, and where it can send data.7. An ethical constraint that can block, with a named human owner, and a record of every decision made under it, including cleared false positives.8. A tested rollback to a prior behavioral bundle.When these minimum elements are present, an organisation can be considered that is following IntDEx even if advanced practices (e.g., advanced cost control, multiple model validation, automated drift detection, full traceability matrices) are introduced gradually. This flexible adoption model allows teams to scale IntDEx according to their governance needs, regulatory environment, and appetite for AI-enabled acceleration, while preserving the core principles of intent-centricity, behavior-driven engineering, and HITL and evidential independence.Relevant templates and examples of files/configurations mentioned in this paper have been produced and will be published but are not part of this paper.## 10 Conclusion### 10.1 Summary of ContributionsThis paper introduced IntDEx, an AI-native digital solutions engineering framework designed to support the development of digital solutions in environments where AI is the main tool for design, coding, documentation, and maintenance. It distinguishes between the AI delivery engine used by the team to build and deliver, and the digital solution delivered to users. IntDEx was developed in response to the practical needs of research, where rapid development for experimentation and adaptive refinement were essential. IntDEx defines a complete AI-native delivery flow consisting of defining characteristics, phases, roles and responsibilities and practices.**External sources**The framework synthesises established methodologies, recent AI-native engineering research and results and observations from practical application (see references).**Use of AI**For the development of the content of the paper AI (MS copilot, ChatGPT, Gemini) has been used for grammar and style checking, improving consistency and readability, as a search engine for relevant open resources. The author is sole accountable for this work.### 10.2 Implications for ResearchThe introduction of IntDEx contributes to the emerging field of AI-native software engineering by providing a methodology grounded in both theory and practice. In some areas IntDEx proposes novel methods and is new as whole, but at the same time it is also ready to be used with minimum or even no business organisational changes. Future research may expand different areas for the framework, provide improved templates and enabling tools, further practical validation, revised practices, business models etc. In addition, advanced automated prompt evaluation, adaptive model routing strategies, and advanced drift detection mechanisms could be explored. Additional studies may examine how IntDEx performs across different domains, organisational contexts, and regulated environments.### 10.3 Implications for PracticeAI-native development represents a paradigm shift in how digital solutions are created. IntDEx can evolve in a robust, academically grounded, and practically validated methodology for navigating this shift or provide ideas practices and concepts to be used but other frameworks and methodologies.The framework supports rapid delivery while maintaining governance, human augmentation, ethical compliance, documentation and risk (including, security, cost, sovereignty) management, making it ideal for digital transformation initiatives that rely on AI in complex environments. IntDEx provides a foundation using AI-native approaches that are safe, reliable, ethical, and aligned with expectations for faster to deliver and better in quality digital solutions, via a modern Delivery Experience.Back to home
Comments
Sign in to add and view your comments and replies.