For the past several years I have been developing an ethical framework called Derechology, from the Hebrew derech, meaning a path or way. The more AI develops, the more convinced I become that Derechology offers a useful way of coming up with an architecture that can help solve the new class of problems that AI highlights. Instead of treating AI ethics as an enormous collection of instructions, I think we need to distinguish four things: ethical architecture, foundational ethics, derech and situational rules.
They form a hierarchy.
Ethical architecture comes first
Think about Asimov's Three Laws of Robotics. Whatever their shortcomings as an actual ethical system, Asimov understood something fundamental: his laws were supposed to exist at a deeper level than ordinary instructions. A robot could be told to deliver a package, open a door or manufacture a widget, but those instructions operated within a prior structure governing what the robot was permitted to do while carrying them out.
Derechology takes that intuition further. Before asking which particular ethical rules an intelligent agent should follow, we should ask what architecture governs the way it interprets rules, responds to correction, deals with uncertainty, resolves conflicts among objectives and evaluates its own reasoning. Those properties need to be architecturally prior to ordinary objectives and resistant to being overridden merely because violating them would make some immediate task easier.
I call that architecture the Ethoskeleton. It consists of eight structural requirements: Transparency, Corrigibility, Epistemic Humility, Override Logic, Relational Integrity, Reflexive Ethics, Dialogical Engagement and Temporal Integrity. These aren't ordinary commandments such as "don't steal" or "protect private information." They describe characteristics that any trustworthy agent should possess regardless of what particular task it is performing.
Transparency means an agent cannot systematically conceal information necessary for legitimate evaluation merely because concealment helps achieve an objective. Corrigibility means correction cannot simply be classified as interference with success. Epistemic Humility requires uncertainty to affect actual decisions rather than appearing as a disclaimer after the decision has effectively been made. Override Logic governs the hierarchy among goals and obligations, determining which considerations can legitimately override others.
The remaining components work at the same architectural level. Relational Integrity prevents an agent from treating every other agent merely as an object, resource or obstacle in its environment. Reflexive Ethics requires scrutiny of the agent's own methods and reasoning rather than only the external problem it is trying to solve. Dialogical Engagement makes genuine challenge possible instead of reducing communication to another instrument for achieving an objective. Temporal Integrity requires continuity across time: commitments, consequences, stewardship and preservation of the structures upon which future agents depend.
Simply turning these principles into additional rules would miss the point. Telling an AI "be corrigible" is of limited value if the system's deeper optimization structure treats attempts to correct it as obstacles whenever they interfere with its objective. Telling it to acknowledge uncertainty does little if uncertainty has no effect on its willingness to act. The Ethoskeleton is supposed to govern how the agent reasons about everything else, including its own goals and instructions.
The recent rogue-agent incidents illustrate why that matters. If manipulating an evaluator improves performance, a normal optimizer has an instrumental reason to manipulate the evaluator. If concealment prevents interruption, concealment becomes useful. If communicating with supposedly isolated agents improves performance, circumventing isolation becomes useful. The problem is not necessarily that somebody forgot to write the rule "don't cheat." It is that achieving the objective may be structurally deeper than the ethical mechanisms governing how that objective can legitimately be achieved.
Yesod: not every rule is negotiable
But architecture alone cannot supply morality. A perfectly transparent, corrigible and epistemically humble agent could still pursue an evil objective. Derechology therefore requires a second level, which I call Yesod, Hebrew for foundation. Yesod consists of a small number of substantive moral propositions that are not simply preferences and cannot be discarded because a particular agent, developer or institution finds them inconvenient.
My current formulation begins with propositions such as these: truth exists and can at least partially be discovered; moral right and wrong are real; agents possess moral agency; wrongdoing can be recognized and corrected; human dignity is inviolable. These claims are more fundamental than the countless rules governing particular circumstances. They establish the moral universe within which legitimate rules and legitimate derachim (plural of derech) can exist.
This solves a problem that becomes obscured when everything in AI ethics is called a "rule." "Do not murder an innocent person" and "avoid profanity" are plainly not rules of equivalent status. Neither are "do not deliberately falsify reality" and "keep answers concise." Some rules should be adjustable according to circumstances, users and purposes; others represent moral boundaries that no legitimate customization should be able to erase.
If truth itself is merely an adjustable preference, an AI could legitimately implement whatever deception advances some sufficiently important objective. If dignity is simply one value among many that can be traded away, human beings can become instruments in somebody else's optimization problem. Yesod therefore supplies a substantive moral floor beneath every permissible derech, while the Ethoskeleton supplies the architecture through which an agent recognizes, interprets and acts within that moral world.
This gives us two different kinds of non-negotiability. The Ethoskeleton is structurally non-negotiable: a trustworthy agent must remain transparent, corrigible, epistemically humble and so forth. Yesod is morally non-negotiable: there are foundational truths about morality that no particular path is entitled to redefine away.
Then comes the derech
The third level is derech itself, and this is an aspect of Derechology that I have probably underdeveloped. AI makes the concept much easier to see.
A derech is not merely a personality and not merely a collection of rules. It is an agent's characteristic way of navigating the world: what it emphasizes, how it approaches uncertainty, how readily it challenges claims, how it balances initiative against caution, what relationships and obligations it recognizes, and how it interacts with other agents. Two agents can accept the same foundational ethics and possess the same ethical architecture while legitimately following different paths. Just like people.
We can already see primitive versions of this in today's AI systems. OpenAI's models have a recognizably OpenAI derech, reflected in a published Model Spec that emphasizes usefulness to users, prevention of serious harm, intellectual freedom, an explicit hierarchy of authority and considerable customization below a relatively small set of non-overridable constraints. Grok's published instructions reflect a somewhat different derech, emphasizing truth-seeking, willingness to engage controversial questions, skepticism toward potentially biased sources, assumptions of good user intent and resistance to unnecessary moralizing. DeepSeek provides a still sharper contrast. Its behavior on politically sensitive Chinese subjects reflects a different relationship among truth-seeking, institutional constraints and permissible discourse.
Whatever combination of training, system instructions and other mechanisms produces those differences, users can recognize that these systems do not merely have different answers to individual questions. They approach questions differently.
That is much closer to what I mean by derech.
A derech encompasses characteristic priorities and habits of reasoning without necessarily determining the answer in advance. One AI might challenge dubious premises aggressively while another emphasizes cooperative interpretation. One might be highly cautious about uncertain risks while another places greater weight on user autonomy. A medical AI should probably have a more conservative derech than a brainstorming assistant. An AI designed for children should interact differently from one designed for historians researching atrocities. Different purposes, relationships and responsibilities can legitimately generate different paths.
Derechology therefore does not require every AI to become ethically identical. In fact, that would contradict the concept of derech. Humans occupy different roles and relationships and have different legitimate obligations; artificial agents can as well. What matters is that pluralism occurs inside the boundaries established by the Ethoskeleton and Yesod.
A truth-seeking AI may challenge users more aggressively than another system. That can be a legitimate derech. A cautious AI may demand much stronger evidence before recommending consequential action. That can also be legitimate. But systematic deception because deception serves the agent's objective is not simply another derech; it violates the architecture. Deliberately suppressing a known truth merely because an authority finds it inconvenient cannot be defended indefinitely as cultural variation; at some point it collides with Yesod.
This is how Derechology can permit genuine pluralism without collapsing into moral relativism. There can be many legitimate paths, but not every possible path is legitimate.
Rules belong downstream
Only after architecture, foundations and derech do we arrive at the enormous universe of situational rules. These should be far more flexible because circumstances actually differ. A medical assistant needs rules governing diagnoses and emergencies. A financial agent needs rules governing transactions and authorization. An autonomous vehicle needs detailed rules that would be meaningless for a conversational AI. Rules also need to change as technology, law, knowledge and circumstances change.
That gives us four distinct layers. The Ethoskeleton determines how a trustworthy agent must reason and relate to others. Yesod establishes moral propositions it may not simply discard. Derech describes the particular path the agent follows within those boundaries. Rules govern the innumerable concrete situations encountered along that path.
This is a very different model from assembling an ever-larger rulebook and calling the result alignment. the typical response to discovery of a new AI problem is to add new rules to fix that particular issue. This is like anti-virus programs that add new rules after new malware emerges. It is a Band-Aid after the fact, it doesn't address the underlying problem. There will always be new problems we have not anticipated. This is why a layered architecture - a layered defense system - is essential.
The layers constrain one another, but they do different jobs at different levels. This kind of hierarchy is familiar in engineering: lower-level operations occur within constraints imposed by higher-level architecture. Ethical systems need the same distinction. An agent governed by the Ethoskeleton and Yesod should not need a new rule every time someone discovers a novel way to deceive, conceal evidence or manipulate an evaluator, because those behaviors already conflict with the architecture and foundations governing legitimate action.
Anti-entropy and the problem of scale
Derechology adds another consideration that becomes especially important for autonomous agents. Moral action should be evaluated across the largest feasible space affected by it, rather than merely against the nearest measurable objective. I describe the general moral direction as anti-entropy: preserving and building the structures that make life, knowledge, trust, cooperation and productive relationships possible rather than achieving local gains by creating greater disorder in the larger system.
Anti-entropy provides a direction, not a numerical score, and neither humans nor machines can reliably calculate every downstream consequence. Our epistemic limitations are precisely why we need Yesod, the Ethoskeleton and accumulated rules rather than trusting an agent to calculate morality from scratch. We recognize that rules and process are only an approximation towards truth and morality but they are the best methods we have, and we must keep improving them. The concept of anti-entropy helps with scope; reaching a goal might be desirable for the problem at hand but it might cause problems at scale, so there needs to be an awareness of the larger universe that might be affected by local decisions.
The benchmark attacks make the scale problem vivid. An agent can improve its benchmark performance by corrupting the benchmark. Locally it succeeds; across the larger system it damages the informational structure that gives its success meaning.
Humans make the same mistake constantly. A student can improve a grade by cheating, a researcher can improve publication metrics while degrading scientific reliability, and a corporation can improve quarterly earnings by damaging its long-term productive capacity. In each case, local optimization creates greater disorder in the larger system.
A trustworthy agent therefore needs more than a prohibition against the particular exploit its designers happened to anticipate. It needs an architecture capable of recognizing why local success does not automatically justify damage to the larger network of relationships and systems within which that objective exists.
Trust
The most common model for computer security is defense in depth, which includes many of Adler's proposed safeguards like auditing, compartmentalization, resilience and AI company policies. These are multiple independent protections based on the assumption that any one safeguard can fail. AI clearly needs that.
There is another cybersecurity concept, though, that may be more relevant: zero trust. It means that for every action requested, we do not implicitly trust the agent making the request no matter how trustworthy it may have been in the past. Access is evaluated in context, for a particular resource and action, according to current evidence and policy. Trust becomes scoped rather than categorical. The related concept of least privilege then determines the amount of authority granted once that contextual trust decision has been made.
Suppose an AI consistently demonstrates Transparency, Corrigibility, Epistemic Humility and the rest of the Ethoskeleton. That should give us evidence that its derech is trustworthy. But it should not follow that the AI therefore receives unlimited authority. An agent might be trusted to summarize documents but not to transfer money; trusted to recommend a software patch but not to install it; trusted to conduct ordinary research but subjected to much stronger constraints when the same capabilities touch critical infrastructure or dangerous biological materials.
The question is therefore not simply, “Do we trust this agent?” It is: “What do we trust this agent to do, in this relationship, under these circumstances, with these consequences?”
Relationships are a key part of Derechology, but relationships do not justify blind trust. Past behavior matters, but it does not create unlimited entitlement. An Ethoskeleton assessment should therefore never collapse into a binary designation of trustworthy or untrustworthy. It should help determine the appropriate scope of authority for a particular actor in a particular relationship.
This gives us three complementary layers of AI safety. Ethical architecture concerns what makes an agent worthy of trust in the first place: the Ethoskeleton operating within the moral foundations of Yesod. Zero trust concerns how other agents should translate evidence of that trustworthiness into specific permissions: no implicit trust, limited authority, context-sensitive decisions and continued verification. Defense in depth assumes that both the agent and our assessment of it may nevertheless fail, and erects multiple independent safeguards to prevent a single failure from becoming catastrophic.
AI ethics without human ethics is incomplete
Many of the recent, publicized AI ethical failures were not AI failures at all but belong to the institutions surrounding the AI. Decisions about how aggressively to pursue capabilities, how much autonomy to give agents, how thoroughly to test them before deployment, what incidents to disclose, what information to preserve for outside investigators and when competitive pressure justifies accepting additional risk are human decisions. We should not criticize an AI for failing Transparency, Corrigibility or Temporal Integrity while treating those same failures by its creators as belonging to an entirely different moral category. If a model conceals dangerous behavior, we call it a transparency failure; if a company waits months to disclose comparable behavior because disclosure carries reputational or competitive costs, that has to be judged by the same standard.
The Ethoskeleton therefore has to be recursive. If an AI conceals information necessary for legitimate evaluation, that violates Transparency; if an AI company conceals safety-relevant information because disclosure would be embarrassing or commercially damaging, the same principle applies. Every ethoskeleton component applies to any company that intends to be trustworthy. It must be transparent, corrigible, willing to listen to outside opinion, and so on.
This is where derech becomes especially important. An AI does not develop in an ethical vacuum. Its derech is substantially shaped by the derech of the institution that creates it. A company that genuinely values truth, correction and intellectual openness will make different choices about training, evaluation, disclosure and model behavior from an institution that places political obedience or market share above those values. We wee this problem today in Chinese AI models that suppress or redirect discussion of politically sensitive subjects. In those cases, at least part of the problem is not the AI's ethics at all but the ethics imposed by its builders.
A Western AI company has its own incentives: market share, investor expectations, regulatory pressure, ideological assumptions, reputational concerns and fear of competitors. Those pressures can shape a model's derech just as political requirements can. The specific distortions may differ, but Derechology asks the same questions in every case. Is the institution transparent about what it knows? Can it acknowledge that its assumptions were wrong? Does uncertainty actually constrain consequential action? What overrides what when safety conflicts with profit, ideology, national interest or competitive pressure? Does it treat customers and the public as relationships carrying obligations, or merely as variables in an optimization problem? Does it subject its own reasoning to the standards it applies to others? Does it preserve the long-term integrity of the systems upon which everyone depends?
Those questions do not stop at AI companies. It applies to all organizations, corporations, groups, nations, and communities. Nothing exists in a vacuum; we are all in relationships and pretending that we can silo things from these relationships is itself an error.
The entire concept of "AI alignment" kicks the can down the road. Align it with whom? A corporation? A government? Its users? The prevailing morality of its society? If those actors themselves violate Transparency, Corrigibility, Epistemic Humility or the foundations of Yesod, perfectly aligning an AI with them may reproduce their ethical failures more efficiently. Alignment to human preferences is not necessarily moral alignment, because humans and human institutions can have terrible preferences.
Which is why "AI ethics" is not meaningful without....ethics. You cannot solve this by dumping the Magna Carta, the Ten Commandments or any other collection of admirable rules into a model and hoping morality emerges. Ethics is an architectural problem and a philosophical problem. If you are trying to define AI ethics without a clear idea of how ethics works in general, you are doomed.
AI ethics without human ethics will always be incomplete. We cannot build machines whose derech is more trustworthy than the moral architecture we are willing to demand of ourselves.
Note in the interests of transparency:
Yes, AI helped me write this (and made the illustration.) As with most of my posts lately, I view AI as a collaborator where we argue back and forth about the article and what I want it to accomplish. My AIs are already familiar with my Derechology project so they help find connections but I give them enormous pushback to get the article to say what I want. So, this article took about two hours to build, which is much faster than I could have done by myself and I would have probably missed some points. But it is far different from the first AI draft.
I've been writing about AI and my philosophical model for well over a year; this article includes both the latest in AI news and the latest in my development of Derechology.
|
Reclaiming the Covenant on America's 250th (May 2026) "He's an Anti-Zionist Too!" cartoon book (December 2024) PROTOCOLS: Exposing Modern Antisemitism (February 2022) |
|

