Metadata-Version: 2.4
Name: opteryx_access
Version: 0.2.12
Summary: Permission checks and grant/revoke for the Opteryx platform
Author-email: joocer <justin.joyce@joocer.com>
License:                                  Apache License
                                   Version 2.0, January 2004
                                http://www.apache.org/licenses/
        
           TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
        
           1. Definitions.
        
              "License" shall mean the terms and conditions for use, reproduction,
              and distribution as defined by Sections 1 through 9 of this document.
        
              "Licensor" shall mean the copyright owner or entity authorized by
              the copyright owner that is granting the License.
        
              "Legal Entity" shall mean the union of the acting entity and all
              other entities that control, are controlled by, or are under common
              control with that entity. For the purposes of this definition,
              "control" means (i) the power, direct or indirect, to cause the
              direction or management of such entity, whether by contract or
              otherwise, or (ii) ownership of fifty percent (50%) or more of the
              outstanding shares, or (iii) beneficial ownership of such entity.
        
              "You" (or "Your") shall mean an individual or Legal Entity
              exercising permissions granted by this License.
        
              "Source" form shall mean the preferred form for making modifications,
              including but not limited to software source code, documentation
              source, and configuration files.
        
              "Object" form shall mean any form resulting from mechanical
              transformation or translation of a Source form, including but
              not limited to compiled object code, generated documentation,
              and conversions to other media types.
        
              "Work" shall mean the work of authorship, whether in Source or
              Object form, made available under the License, as indicated by a
              copyright notice that is included in or attached to the work
              (an example is provided in the Appendix below).
        
              "Derivative Works" shall mean any work, whether in Source or Object
              form, that is based on (or derived from) the Work and for which the
              editorial revisions, annotations, elaborations, or other modifications
              represent, as a whole, an original work of authorship. For the purposes
              of this License, Derivative Works shall not include works that remain
              separable from, or merely link (or bind by name) to the interfaces of,
              the Work and Derivative Works thereof.
        
              "Contribution" shall mean any work of authorship, including
              the original version of the Work and any modifications or additions
              to that Work or Derivative Works thereof, that is intentionally
              submitted to Licensor for inclusion in the Work by the copyright owner
              or by an individual or Legal Entity authorized to submit on behalf of
              the copyright owner. For the purposes of this definition, "submitted"
              means any form of electronic, verbal, or written communication sent
              to the Licensor or its representatives, including but not limited to
              communication on electronic mailing lists, source code control systems,
              and issue tracking systems that are managed by, or on behalf of, the
              Licensor for the purpose of discussing and improving the Work, but
              excluding communication that is conspicuously marked or otherwise
              designated in writing by the copyright owner as "Not a Contribution."
        
              "Contributor" shall mean Licensor and any individual or Legal Entity
              on behalf of whom a Contribution has been received by Licensor and
              subsequently incorporated within the Work.
        
           2. Grant of Copyright License. Subject to the terms and conditions of
              this License, each Contributor hereby grants to You a perpetual,
              worldwide, non-exclusive, no-charge, royalty-free, irrevocable
              copyright license to reproduce, prepare Derivative Works of,
              publicly display, publicly perform, sublicense, and distribute the
              Work and such Derivative Works in Source or Object form.
        
           3. Grant of Patent License. Subject to the terms and conditions of
              this License, each Contributor hereby grants to You a perpetual,
              worldwide, non-exclusive, no-charge, royalty-free, irrevocable
              (except as stated in this section) patent license to make, have made,
              use, offer to sell, sell, import, and otherwise transfer the Work,
              where such license applies only to those patent claims licensable
              by such Contributor that are necessarily infringed by their
              Contribution(s) alone or by combination of their Contribution(s)
              with the Work to which such Contribution(s) was submitted. If You
              institute patent litigation against any entity (including a
              cross-claim or counterclaim in a lawsuit) alleging that the Work
              or a Contribution incorporated within the Work constitutes direct
              or contributory patent infringement, then any patent licenses
              granted to You under this License for that Work shall terminate
              as of the date such litigation is filed.
        
           4. Redistribution. You may reproduce and distribute copies of the
              Work or Derivative Works thereof in any medium, with or without
              modifications, and in Source or Object form, provided that You
              meet the following conditions:
        
              (a) You must give any other recipients of the Work or
                  Derivative Works a copy of this License; and
        
              (b) You must cause any modified files to carry prominent notices
                  stating that You changed the files; and
        
              (c) You must retain, in the Source form of any Derivative Works
                  that You distribute, all copyright, patent, trademark, and
                  attribution notices from the Source form of the Work,
                  excluding those notices that do not pertain to any part of
                  the Derivative Works; and
        
              (d) If the Work includes a "NOTICE" text file as part of its
                  distribution, then any Derivative Works that You distribute must
                  include a readable copy of the attribution notices contained
                  within such NOTICE file, excluding those notices that do not
                  pertain to any part of the Derivative Works, in at least one
                  of the following places: within a NOTICE text file distributed
                  as part of the Derivative Works; within the Source form or
                  documentation, if provided along with the Derivative Works; or,
                  within a display generated by the Derivative Works, if and
                  wherever such third-party notices normally appear. The contents
                  of the NOTICE file are for informational purposes only and
                  do not modify the License. You may add Your own attribution
                  notices within Derivative Works that You distribute, alongside
                  or as an addendum to the NOTICE text from the Work, provided
                  that such additional attribution notices cannot be construed
                  as modifying the License.
        
              You may add Your own copyright statement to Your modifications and
              may provide additional or different license terms and conditions
              for use, reproduction, or distribution of Your modifications, or
              for any such Derivative Works as a whole, provided Your use,
              reproduction, and distribution of the Work otherwise complies with
              the conditions stated in this License.
        
           5. Submission of Contributions. Unless You explicitly state otherwise,
              any Contribution intentionally submitted for inclusion in the Work
              by You to the Licensor shall be under the terms and conditions of
              this License, without any additional terms or conditions.
              Notwithstanding the above, nothing herein shall supersede or modify
              the terms of any separate license agreement you may have executed
              with Licensor regarding such Contributions.
        
           6. Trademarks. This License does not grant permission to use the trade
              names, trademarks, service marks, or product names of the Licensor,
              except as required for reasonable and customary use in describing the
              origin of the Work and reproducing the content of the NOTICE file.
        
           7. Disclaimer of Warranty. Unless required by applicable law or
              agreed to in writing, Licensor provides the Work (and each
              Contributor provides its Contributions) on an "AS IS" BASIS,
              WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
              implied, including, without limitation, any warranties or conditions
              of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
              PARTICULAR PURPOSE. You are solely responsible for determining the
              appropriateness of using or redistributing the Work and assume any
              risks associated with Your exercise of permissions under this License.
        
           8. Limitation of Liability. In no event and under no legal theory,
              whether in tort (including negligence), contract, or otherwise,
              unless required by applicable law (such as deliberate and grossly
              negligent acts) or agreed to in writing, shall any Contributor be
              liable to You for damages, including any direct, indirect, special,
              incidental, or consequential damages of any character arising as a
              result of this License or out of the use or inability to use the
              Work (including but not limited to damages for loss of goodwill,
              work stoppage, computer failure or malfunction, or any and all
              other commercial damages or losses), even if such Contributor
              has been advised of the possibility of such damages.
        
           9. Accepting Warranty or Additional Liability. While redistributing
              the Work or Derivative Works thereof, You may choose to offer,
              and charge a fee for, acceptance of support, warranty, indemnity,
              or other liability obligations and/or rights consistent with this
              License. However, in accepting such obligations, You may act only
              on Your own behalf and on Your sole responsibility, not on behalf
              of any other Contributor, and only if You agree to indemnify,
              defend, and hold each Contributor harmless for any liability
              incurred by, or claims asserted against, such Contributor by reason
              of your accepting any such warranty or additional liability.
        
           END OF TERMS AND CONDITIONS
        
           APPENDIX: How to apply the Apache License to your work.
        
              To apply the Apache License to your work, attach the following
              boilerplate notice, with the fields enclosed by brackets "[]"
              replaced with your own identifying information. (Don't include
              the brackets!)  The text should be enclosed in the appropriate
              comment syntax for the file format. We also recommend that a
              file or class name and description of purpose be included on the
              same "printed page" as the copyright notice for easier
              identification within third-party archives.
        
           Copyright 2026 Justin Joyce (@joocer)
        
           Licensed under the Apache License, Version 2.0 (the "License");
           you may not use this file except in compliance with the License.
           You may obtain a copy of the License at
        
               http://www.apache.org/licenses/LICENSE-2.0
        
           Unless required by applicable law or agreed to in writing, software
           distributed under the License is distributed on an "AS IS" BASIS,
           WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
           See the License for the specific language governing permissions and
           limitations under the License.
        
Project-URL: Homepage, https://github.com/mabel-dev/opteryx-access
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Operating System :: OS Independent
Requires-Python: >=3.11
Description-Content-Type: text/markdown
License-File: LICENSE
Provides-Extra: firestore
Requires-Dist: google-cloud-firestore==2.*; extra == "firestore"
Provides-Extra: dev
Requires-Dist: pytest>=8; extra == "dev"
Requires-Dist: pytest-cov>=5; extra == "dev"
Requires-Dist: opteryx_access[firestore]; extra == "dev"
Dynamic: license-file

# opteryx-access

Permission checks and grant/revoke for the Opteryx platform, as an installable
library rather than a service. Every consumer imports `opteryx_access` and
calls it in-process against policies it already has in hand -- a JWT's
`policies` claim, or a `PolicyStore` backed by whatever it uses for storage
(Firestore, today). There is no `opteryx-access` HTTP surface, and this
package makes no network calls of its own beyond what a storage adapter does.

## Scope: data access only

This package models exactly one thing: who can do what to *data* --
`owner`/`writer`/`reader` on workspaces, collections, and datasets. It has no
notion of billing-account roles (`admin`/`member`, minted separately on a
billing account and unrelated to data roles) and enforces no billing checks.

The one place these two systems touch is **workspace genesis**: creating a
workspace is a billing-gated action (the creator must already be a billing
admin -- checked by whatever service owns billing, not this one), and its
result is a data-permission fact -- the creator becomes the new workspace's
`owner`. That handoff is the entire intersection. `opteryx_access.grants.bootstrap_workspace`
implements the data-permission half of it (an explicit list of
`(principal, role)` pairs, `owner` among them) and knows nothing about
billing; the billing gate is the calling service's job, enforced before
`bootstrap_workspace` is ever reached.

## Why this exists

Three independent, subtly incompatible implementations of "does this role
satisfy this requirement" already exist in the fleet:

- **policy.opteryx** / **control.opteryx** (`app/routes/v1/access.py`,
  `app/models/policy.py` -- byte-for-byte duplicated between the two repos):
  a rank-based `ROLES` tuple used to decide who may create/update/revoke a
  policy, and whether a new grant is redundant against one the principal
  already holds.
- **opteryx-core** (`opteryx/managers/permissions/__init__.py`): a
  set-based `ACTION_MAP` deciding which roles may `READ`/`DELETE`/`DROP`/etc.
  a resource once a query actually runs.
- **odata.opteryx** (`app/auth/permissions.py`): a binary
  `role_allows_read` used for listing/visibility, plus its own pattern
  matcher -- `fnmatchcase` where the others use plain `fnmatch`, a difference
  that is latent today (see "Behavior changes") but means the three do not
  state the same rule even where they currently agree.

On top of that, `policy.opteryx`/`control.opteryx`/`odata.opteryx`/
`register.opteryx` each carry their own copy of "parse the `policies` claim
out of a decoded JWT." Comments in several of these point at an
`authorize.opteryx` service (`app.routes.v1.evaluate`) as the semantics every
copy is meant to mirror -- but no such repo exists anywhere in this
workspace. Whether it's a real service in another org/remote or was never
built, every consumer today reimplements its own understanding of "role +
pattern -> allowed" independently, which is exactly the drift this package
is meant to stop.

## What lives here

| Module | Ported from | Purpose |
|---|---|---|
| `roles.py` | `policy.opteryx/app/models/policy.py` | The canonical `ROLES = ("owner", "writer", "reader")` and `role_outranks_or_equals`, the privilege ordering used only for redundancy detection. |
| `actions.py` | `opteryx-core/opteryx/managers/permissions/__init__.py` | `ACTION_ROLES`: which roles may perform `READ`/`WRITE`/`DELETE`/`CREATE`/`DROP`/`ALTER`/`REFRESH`/`MANIFEST`, plus `GRANT`/`REVOKE` (new -- makes policy-administration authority explicit in the same table instead of an implicit rule elsewhere) and `AUTOMATE` (new -- standing automation is the owner's to create; see "Roles" below). |
| `patterns.py` | `policy.opteryx/app/models/policy.py` + `app/routes/v1/access.py` | What a pattern and a principal may look like, and how a pattern matches: `validate_pattern`, `validate_principal`, `resource_matches` (see "Patterns and principals" below). |
| `models.py` | `authenticate.opteryx/app/policies.py` | `Grant` (role+pattern, the JWT-carried shape), `Policy` (principal+role+pattern+metadata, the stored shape) and `Entitlement` (actions+pattern, never stored), plus `parse_policy_claim` for the `[role, pattern]` pairs a token carries. |
| `checks.py` | `opteryx-core`'s `can_perform_action`/`can_perform_workspace_action` + `policy.opteryx`'s `_check_pattern_access`/`_check_workspace_access` | The evaluation layer: data-plane checks over `Grant`s, administrative-plane checks over `Policy` documents. |
| `grants.py` | `policy.opteryx/app/routes/v1/access.py`'s `create_policy`/`update_policy`/`delete_policy`/`create_genesis_policies` | The write half: `grant()`/`update_grant()`/`revoke()`/`revoke_grant()`/`bootstrap_workspace()`, enforcing every rule those routes did (self-grant prevention, pattern authority, conflict detection, principal and pattern validation) before calling a store. `revoke_grant` is the by-value form behind SQL `REVOKE` -- strictly 1:1 resolution by (principal, pattern, role), with level-mismatch diagnostics. Also the two reads that need a store: `grants_for_principal()` and `owned_by()`. |
| `entitlements.py` | (new) | Scoped entitlements: `automation_admin::acme` resolves to `AUTOMATE` on `acme.*`. `ENTITLEMENT_KINDS` declares what each kind confers, the scope is in the name at issue, and none may confer `GRANT`/`REVOKE`. See "Entitlements" below. |
| `store.py` | (new) | The `PolicyStore` protocol: the storage contract, and nothing else. No rules live here -- see "Layering" below. |
| `audit.py` | `policy.opteryx/app/routes/v1/access.py`'s `_audit_policy_change` | One structured record per policy change, on the same field contract the existing log transforms already parse -- see "Audit records" below. |
| `capability.py` | (new) | The permissions capability opteryx-core registers -- the one module that knows the engine exists. See "The opteryx-core capability" below. |
| `adapters/firestore.py` | (new) | `FirestorePolicyStore`, matching the `{workspace}/$policies/access` layout policy.opteryx/control.opteryx already write to -- a drop-in for their inline Firestore calls. |
| `exceptions.py` | (new) | Plain exceptions (`SelfAccessError`, `AccessDeniedError`, `PolicyConflictError`, `InvalidActionError`, ...) instead of `HTTPException` -- each caller translates to its own transport. |

## Layering

Rules and storage are separate, and the dependency only points one way:

```
roles / actions / patterns / models     the vocabulary  (no I/O, no deps)
                 |
              checks.py                 read-side rules (no store at all)
                 |
              grants.py                 write-side rules -- the only module
                 |                       that mutates a store
              store.py                  the storage contract (a Protocol)
                 |
       adapters/firestore.py            one backend
```

Two consequences worth relying on:

- **`checks.py` never touches a store.** It answers from grants handed to it,
  so opteryx-core can evaluate a query's permissions from a JWT with no
  storage, no credentials, and no network.
- **A store enforces nothing.** It translates to its backend and no more. A
  store that filtered out policies it thought invalid, or refused a write it
  thought unauthorized, would be enforcing policy somewhere none of the tests
  for those rules can see it. Every rule lives in `grants.py`, tested against
  an in-memory fake.

## Roles: reader, writer, owner

Three roles, and the line between each pair is a sentence:

- **A reader uses the data.** This is the role for the masses -- everyone
  consuming a data product. It confers `SELECT` and `EXPLAIN` and nothing
  else: no creating, no changing, no removing, nothing that outlives the
  query.
- **A writer changes what is in a relation.** Rows, and the artefacts that
  hold or derive them: tables, views, comments, compaction. Writer work is
  attended and finite -- it is over when the statement finishes -- and what
  it makes can be recreated from its own text.
- **An owner changes what a relation is to everyone else.** Its shape, its
  name, its existence, its physical layout, who may read it, and what it does
  on its own. Each of those alters the contract a reader depends on, or the
  set of readers, so each is the owner's.

`ACTION_ROLES` states this per action:

| Action | Roles | What it covers |
|---|---|---|
| `READ` | reader, writer, owner | `SELECT`, `EXPLAIN`, listing existence and shape |
| `WRITE`, `UPDATE`, `DELETE` | writer, owner | `INSERT`, `MERGE`, `TRUNCATE`, `OPTIMIZE`, `COMMENT ON`, redefining a view |
| `CREATE` | writer, owner | New tables, collections, and views |
| `REFRESH` | writer, owner | Rebuilding a materialized view from its stored definition |
| `DROP` | owner | Removing a table, collection, or workspace -- destroys history |
| `ALTER` | owner | Columns, types, renames, relationships, clustering, snapshots |
| `MANIFEST` | owner | `SHOW MANIFEST FOR` -- file paths and storage layout |
| `AUTOMATE` | owner | Tasks, triggers, and materialized views -- see below |
| `GRANT`, `REVOKE` | owner | Who else holds a role here |

### Automation is owner-tier

`AUTOMATE` gates creating, dropping, suspending, resuming, and re-pinning the
identity of a **task** or **trigger**, and creating, suspending, resuming, or
re-pinning the owner of a **materialized view** (creating one lands a refresh
trigger on every source it reads, and its refreshes run as its owner). It is
the one action whose placement is not obvious from "does this change rows",
so the reasoning is worth stating.

An `INSERT` is over when it finishes. A trigger is a standing commitment: it
runs unattended, indefinitely, as a pinned identity, on the owner's compute,
and it can write to other relations and fire further triggers. That is a
decision about what the relation *does* to the world, not what is in it --
much closer to `GRANT` than to `WRITE`. The person accountable for the
relation is the person who should be asked before it starts acting on its
own.

Every engine with triggers agrees. Postgres and MySQL have a separate
`TRIGGER` privilege that `INSERT`/`UPDATE` do not include; Snowflake gates
tasks on `EXECUTE TASK`, an account-level grant handed out sparingly; SQL
Server requires `ALTER` on the table. None of them treat "make this run by
itself" as an ordinary write.

A writer who wants derived data creates a plain view. Turning that into
something that refreshes itself is precisely the moment an owner should be
consulted, which is why `CREATE MATERIALIZED VIEW` is `AUTOMATE` rather than
`CREATE`. `REFRESH` of an existing materialized view stays writer-tier: the
decision to have it was taken, and authorized, when it was created. `DROP
MATERIALIZED VIEW` stays at `DROP`: it destroys the backing table's history,
which is what `DROP` is for, and both are owner-tier.

### What the information schema shows to whom

`information_schema` is filtered per row by the caller's access to the row's
subject, never gated per table. A user with no grants in a workspace gets
empty results, not an error. Which action gates a row follows from the tiers
above: a row is shown to whoever could act on what it describes.

| Table | Shown at | Why |
|---|---|---|
| `tables`, `columns`, `schemata` | `READ` | Existence and shape are what a reader needs to write a query |
| `column_relationships` | `READ` on **both** ends | A row names a second dataset; half-visible rows leak the far side |
| `views` | `READ` for the row; `WRITE` for `view_definition` | The SQL names relations the reader may hold no grant on, and is a writer's to author |
| `triggers` | `AUTOMATE` on the source table | Only an owner could have made one or can act on it |
| `tasks` | `AUTOMATE` on the task | Nobody `SELECT`s from a task; its statement is automation |

Two things follow from this that are easy to get wrong:

- **`SHOW CREATE` moves with the listing.** A definition hidden in
  `information_schema.views` and freely available through `SHOW CREATE VIEW`
  is Postgres's `pg_catalog` side door: careful filtering in one place, the
  same text one statement away. `SHOW CREATE VIEW` is `WRITE`, `SHOW CREATE
  TASK` is `AUTOMATE`, `SHOW CREATE TABLE` stays `READ` (it is the column
  list and clustering, not the manifest).
- **The tier is per row, not per table.** Roles are held per pattern, so one
  person is owner of some datasets and reader of others in the same
  workspace. A table-level gate ("only owners may query `tasks`") has no
  sensible meaning for them; the row-level one does.

### What opteryx-core asks

The engine names the action; this library decides what it confers. The
binder and `information_schema` ask about these tiers as follows (a
capability must answer `AUTOMATE` like any other action -- a deployment
upgrading this library and opteryx-core does so in step):

- `CREATE`/`DROP` `TASK`, `CREATE`/`DROP`/`ALTER ... TRIGGER` (suspend,
  resume, owner transfer), `CREATE MATERIALIZED VIEW`, and `ALTER
  MATERIALIZED VIEW` (suspend, resume, owner transfer) ask `AUTOMATE`.
  `CREATE TASK ... ON <table>` asks it on the table too, since it lands a
  trigger there.
- `DROP VIEW` asks `WRITE`, matching `CREATE VIEW` and `ALTER VIEW`. A view
  is text and is recreatable, which is the reason `DROP` is owner-only for
  tables and not a reason here.
- `information_schema.triggers` and `.tasks` show a row only where
  `AUTOMATE` holds; `information_schema.views` shows the row at `READ` and
  nulls `view_definition` unless `WRITE` holds.
- `SHOW CREATE VIEW` and `SHOW CREATE MATERIALIZED VIEW` ask `WRITE`, `SHOW
  CREATE TASK` asks `AUTOMATE`, `SHOW CREATE TABLE` asks `READ`.

`SHOW GRANTS` advertises `AUTOMATE` on owner rows, derived from
`ACTION_ROLES` as every action is, so what is advertised and what the engine
asks for moved together.

## Entitlements: operations, not data

A role measures **depth** of access to rows: reader, then writer, then owner,
each strictly containing the one before. That ordering is load-bearing --
`role_outranks_or_equals` uses it and `find_conflict` decides redundancy with
it -- and some authority does not lie on that line at all.

**Running the operations of a workspace is not a deeper form of access to its
data.** Someone who creates, suspends and re-pins its tasks need not read a
row of it; an owner who can read everything is not thereby the right person to
run its automation. So this is a different kind of object -- a set of actions
on a pattern -- and not a fourth role. An `automator` holding `AUTOMATE` but
not `WRITE` neither outranks `writer` nor is outranked by it; putting it in
`ROLES` would mean scoring on the privilege ladder something with no place on
it, and `find_conflict` would start giving wrong answers about redundancy.

### The shape of a name

The identity system already issues named entitlements to accounts --
`platform_admin`, `data_admin`, `user_admin`. An entitlement this package
recognizes is a **kind** and a **scope**, joined by `::`:

```
automation_admin::public
automation_admin::platform
automation_admin::acme
automation_admin::acme.pipelines
```

The kind says what it confers; the scope says where, and covers everything
beneath it. `automation_admin::acme` is `AUTOMATE` on `acme.*`: its holder can
run every task and trigger in `acme` without being able to read a row of it.

```python
ENTITLEMENT_KINDS = {
    "automation_admin": {"AUTOMATE"},
}
```

**The scope is in the name, decided when the entitlement is issued.** The
situation this exists for -- a workspace whose data belongs to a pipeline, run
by bots, with no human owner, whose operations still need a human -- is not
special to the platform's own namespaces. It is any customer whose workspace
works that way, and the platform, not this package, is where "which
workspace" is known. A table with `public.*` and `platform.*` hardcoded here
would have answered the first two instances and needed a code change for
every one after. This answers the general case; `public` and `platform` are
its first two uses.

The price is that the identity system carries a workspace name inside an
entitlement string. Accepted: it is a runtime decision by the platform, and
a runtime decision belongs in the system that makes them, not in a table
that changes by pull request.

### Names in, permissions out

- **Assignment stays where it already is.** Nothing here creates, stores or
  revokes an entitlement. There is already a system that assigns these and
  audits doing so; a second mechanism would mean two places to look when
  answering "why can this person do that".
- **A token carries a name, never a permission.** A token can name a scope,
  because that is the platform's to decide; it cannot name an action, because
  that is not. A forged or stale token can claim `automation_admin::acme`; it
  cannot claim `GRANT` on anything.
- **Unrecognized kinds are ignored, not rejected.** A session holding
  `user_admin` gets nothing here and no error -- that name is another
  service's to interpret.
- **A recognized kind with an unusable scope is skipped in a check and
  rejected at issue.** `automation_admin` with no scope, or
  `automation_admin::$internal`, confers nothing -- never something somewhere
  by default. `resolve_entitlements` skips it, because a permission check is
  not the place to fail a query over a misconfigured account;
  `parse_entitlement_name` raises on it, so the path that mints entitlements
  can refuse to.

### Two namespaces with no administrator

Why this exists at all. `public.*`: no policy can confer anything there
(`validate_pattern` refuses a reserved workspace), the implicit grant caps
everyone at `reader` before issued grants are consulted, and even the
platform identities hold only `writer`. So `AUTOMATE` there had **no holder at
all** -- the platform's own tasks and triggers could not be created,
suspended, re-pinned or dropped by anybody, and
`information_schema.tasks`/`.triggers`, which show a row only where
`AUTOMATE` holds, showed nothing to anyone.

`platform.*` is the same need from the opposite direction: an ordinary
workspace where policies are perfectly legal, but which is run by bots, so no
human holds owner. And a customer's `acme.*` is the same again.

### Consulted before everything else

| | Policy | Entitlement |
|---|---|---|
| Confers | a role | a set of actions |
| Assigned by | `grant()`, into a `PolicyStore` | the identity system, as a name |
| Reserved workspaces | refused | reached |
| Carried in a token as | `policies` (role + pattern) | `entitlements` (names) |
| Scope decided | in the stored policy | in the name, at issue |
| Actions decided | by `ACTION_ROLES`, from the role | by `ENTITLEMENT_KINDS`, from the kind |

Entitlements are evaluated ahead of both the implicit-grant cap and the issued
grants. Above the cap, because `AUTOMATE` on `public.*` is otherwise reachable
by nobody. Above issued grants, because a workspace run by bots has no human
owner to grant it.

**On an owned workspace this deliberately reaches past the owner.** That is
the intent, not a leak: it is the platform saying who runs the operations of a
namespace, a decision above any one workspace's owners, issued by the
platform's identity system, which is the platform's to run. `ENTITLEMENT_KINDS`
stays a short, code-declared table for the same reason -- a *kind* is a new
sort of authority and should change by review; a *scope* is one more instance
of an existing sort and need not.

**It is additive, never subtractive.** `automation_admin::acme` does not make
its holder a writer in `acme`; every other action there is still decided by
that workspace's policies.

**It can never confer policy administration.** `GRANT` and `REVOKE` are
excluded from `ENTITLEABLE_ACTIONS`, checked when the kinds are declared (at
import) and again in `entitlement_permits`. An entitlement is authority
granted outside the ownership model; letting it confer authority *over* that
model would let a non-owner mint ownership and make every other check here
advisory. Engine-private (`$`) names are refused on the same basis, and
`information_schema` cannot be a scope for the same reason it cannot be a
pattern.

### Where they show up

`SHOW GRANTS` lists them first, matching the order `can_perform_action`
decides in. The `role` column reads `entitlement` -- deliberately not one of
`ROLES`, so the row cannot be mistaken for a role that was granted:

```
pattern           level      role         actions
public.*          workspace  entitlement  AUTOMATE
platform.*        workspace  entitlement  AUTOMATE
personal.alice.*  workspace  owner        ALTER, AUTOMATE, CREATE, ...
public.*          workspace  reader       READ
```

The engine passes the session's entitlement names to `grants()` as a third
argument. It is optional, so an engine that does not pass them gets exactly
the listing it got before entitlements existed -- but a deployment whose
tokens carry names and whose `SHOW GRANTS` omits them reports less than it
enforces.

`SHOW USER` lists every entitlement an account holds, flat, with no signal
which of them mean anything to data access. `entitlement_kinds()` returns the
kinds this package acts on and `parse_entitlement_name()` says whether a
given name is one of them, so that listing can mark the rows that do.

## Patterns and principals

A policy says **who** gets **what role** over **which resources**.

*Which resources* is a pattern: `workspace[.collection[.dataset]]`, where each
dot-separated segment written is either a literal name or `*`.

**A `*` covers everything below it, not one level.** `analytics.*` grants
`analytics.sales` and `analytics.sales.q1` alike -- a grant over a workspace is
a grant over what is in it. Worth stating plainly, because reading `*` as "one
segment" understates what a policy confers.

Rules, all enforced by `validate_pattern`:

- **Names are lowercase**, `a-z0-9_`, starting with a letter. Patterns and
  resource names are normalized before they are compared, so matching is
  case-insensitive and decided identically on every platform.
- **The workspace segment must be literal.** A policy always says which
  workspace it applies to -- `analytics.*` is fine, a bare `*` is not.
- **A segment is either `*` or a whole literal name** -- no partial globs like
  `pub*`. This is what makes the reserved-workspace check a plain membership
  test rather than a match against every reserved name, with no evasion to
  reason about.
- `public`, `personal`, and `information_schema` remain non-grantable.

## Implicit grants, and the `public` exception

Some access is held without a policy having been issued for it, declared once in
`checks.implicit_grants`:

| Who | Holds | Why |
|---|---|---|
| Any identity | `owner` on `personal.<identity>.*` | Your own namespace |
| Everyone | `reader` on `public.*` | Shared open data, readable by all |
| `PLATFORM_IDENTITIES` | `writer` on `public.*` | Something has to load and compact it |

Note the shape: an identity-keyed exception to what a role could confer, which
no policy can be written for. Entitlements generalize it to authority the
identity system issues by name -- see "Entitlements: operations, not data"
above -- and are consulted just ahead of the table below.

These are checked **before** issued grants and **cap** what they cover: a
resource in `public.` or in your own `personal.` is answered there and never
falls through, which is what makes `public.` read-only however broad a policy
someone holds over it.

`public` is where curated open data lives -- GDELT, vulnerability feeds -- and
something has to write it and keep it compacted. Because `public` is a reserved
workspace, `validate_pattern` refuses to store a policy over it, so that access
cannot be issued as an ordinary grant to the identities that do the work. It is
declared instead as a third implicit grant, held by a short closed list of
platform identities (`federator`, `xb500`), ordered ahead of the reader grant it
would otherwise be capped by.

Three things worth being explicit about, since this is a carve-out:

- **Writer, not owner.** These identities load and compact `public`; dropping a
  public dataset or granting anyone access to one is not theirs to do.
- **It is visible.** `SHOW GRANTS` reports implicit grants first, in evaluation
  order, so a platform identity's write access over `public` is legible in the
  same place a policy row would have been.
- **It rests on identity issuance.** Nothing here can verify that a session
  claiming to be `federator` is the platform -- an identity arrives already
  authenticated. These names must be unregisterable wherever accounts are
  created, or the carve-out is a signup form away from anyone.

*Who* is a named individual, enforced by `validate_principal`. There is no
wildcard principal and no group principal: a grant everyone holds is not
something any listing surfaces as unusual, and groups will be their own
concept rather than a pattern smuggled into this field.

Identities are casefolded too, so `XB500` and `xb500` are one principal rather
than two people each holding half the access. Normalizing on write alone would
not be enough -- a lookup for `xb500` would silently miss a stored `XB500` --
so every read that filters by principal normalizes the same way.

## Audit records

Every successful change to a policy emits one structured record, after the
write lands. `grant`, `update_grant`, `revoke`, and `bootstrap_workspace` all
do this themselves -- not the caller -- because a trail assembled by whoever
remembers to assemble it is one refactor away from having a hole in it, and a
hole here is invisible until someone needs it.

Records go to the standard-library logger `opteryx_access.audit` at INFO.
**Being a library, this configures nothing** -- no handlers, no levels. A
service that does not configure logging for that logger will not see these at
all, which for an audit trail is worth confirming rather than assuming.

Each record is emitted twice over, so it survives whatever is in front of it:
as compact single-line JSON in the message (Cloud Run promotes a JSON stdout
line into `jsonPayload` on its own), and via `extra={"json_fields": ...}`
(what google-cloud-logging's handlers read directly). A service that would
rather route these through its own audit channel can call `set_audit_sink`
and be handed the payload dict.

```json
{"event": "policy.updated", "actor": "alice", "workspace": "analytics",
 "policy_id": "9f2c...", "principal": "xb500", "role": "writer",
 "pattern": "analytics.sales.q1", "previous_role": "reader",
 "previous_pattern": "analytics.sales.*",
 "timestamp": "2026-08-12T08:06:37.516307+00:00"}
```

`event` is `policy.created`, `policy.updated`, or `policy.deleted`. The field
names deliberately match what policy.opteryx/control.opteryx already emit, so
the transforms building `opteryx.ops.policy_changes` keep working unchanged
when a service moves onto this package. Absent fields are omitted rather than
null: no `previous_*` on a create, no `role`/`pattern` on a delete. Genesis
grants additionally carry `bootstrap: true`, since they clear none of the
usual authority checks and are worth telling apart from an ordinary grant.

Only changes that actually happened are recorded. A refused grant writes
nothing -- see "Not recorded" at the end of this file for what that leaves
out.

## Which check to call

- **"May this identity administer grants on this pattern?"** ->
  `checks.can_administer_pattern`, over stored `Policy` documents. Requires
  `owner`, and requires that ownership to *cover* the pattern in question.
- **"May this role perform this action on this resource?"** ->
  `checks.can_perform_action`, over `Grant`s. Answered from `ACTION_ROLES`.

They take different inputs because they need different things: a `Grant` is
role plus pattern, all a data action needs; a `Policy` also names the
principal it was issued to, which is what an administrative check reasons
about. Call the one that fits the question -- neither answer substitutes for
the other.

## The opteryx-core capability

opteryx-core allows everything on its own: a CLI or embedded engine has no
workspaces to own and no policy service to have issued anything, so access
control is a property of a deployment rather than of the engine. A deployment
installs this library over that intrinsic default at start-up:

```python
import opteryx
import opteryx_access

opteryx.register_permissions_capability(opteryx_access.capability())
```

From then on the engine's permission gates and its `SHOW GRANTS` are both
answered from here, so what it enforces and what it reports come from one
evaluation.

`opteryx_access.capability` is the **only** module in this package that knows
opteryx-core exists, and even it does not import it -- the engine hands over
an execution context and the adapter reads two attributes off it. Neither
package depends on the other; a deployment brings them together. That is what
keeps opteryx-core's zero-dependency contract intact, and it means the whole
extent of the coupling can be read in one file.

The capability also carries the engine's grant-administration surface --
`apply_grant`, `apply_revoke`, `grants_on`, and `effective_grants_on`,
behind opteryx's `GRANT`/`REVOKE`/`SHOW GRANTS ON`/`SHOW EFFECTIVE GRANTS ON`
statements. `grants_on` lists what is stored AT an object, 1:1 with what a
GRANT or REVOKE there would act on; `effective_grants_on` lists every policy
that COVERS it, so a dataset reachable only through the workspace owner's
`w.*` names that owner instead of returning nothing. Both decide coverage
with `resource_matches`, the matcher that decides real queries, so a listing
cannot report access the engine would not grant. All four are thin
delegations to `grants.py`/`checks.py` (the rules live there, once) and all
four require
the capability to have been built as `capability(store)`; without a store
they raise `PolicyStoreRequiredError` rather than guessing. Because the SQL
surface performs `GRANT`/`REVOKE`, `SHOW GRANTS` reports those actions on
owner rows alongside the data actions -- what is advertised and what the
surface offers moved together.

One thing that stays deliberate:

- **Registration is start-up only.** opteryx-core refuses a capability
  registered after a permission check has already been answered, rather than
  let one process decide the same question two ways.

## Usage

Every check function takes grants/policies already in hand -- it doesn't
fetch them itself. Where those come from depends on what you're holding:

- **A JWT**: `opteryx_access.models.parse_policy_claim(claims)` -- the token
  was already scoped to its own holder when it was minted, so nothing more
  to filter.
- **A `PolicyStore`** (live policy state, not a token): `opteryx_access.grants.grants_for_principal(store, workspace=..., identity=...)`
  -- fetches that identity's issued grants and converts them to the same
  `Grant` shape.

One question is asked across workspaces rather than within one:
`opteryx_access.grants.owned_by(store, identity=...)` returns every policy, in
any workspace, that makes an identity an **owner**. That is what offboarding
needs -- a workspace whose last owner is removed cannot be administered by
anyone, so those grants have to be reassigned before the identity goes. Each
returned `Policy` carries its `workspace`.

There is deliberately no "everywhere this identity can read" equivalent: it
would mostly return reader rows nobody acts on, while being the expensive
query shape and widening what a backend has to index. If a use for it turns
up, it should arrive as its own named operation with that use written down.

```python
from opteryx_access import Grant, can_perform_action

grants = [Grant(role="writer", pattern="analytics.sales.*")]
can_perform_action(grants, "analytics.sales.q1", "DELETE")  # True
can_perform_action(grants, "analytics.sales.q1", "DROP")  # False -- writer, not owner
```

```python
from opteryx_access.grants import grants_for_principal
from opteryx_access.adapters.firestore import FirestorePolicyStore

store = FirestorePolicyStore(db)
grants = grants_for_principal(store, workspace="analytics", identity="bob")
can_perform_action(grants, "analytics.sales.q1", "DELETE")
```

```python
from opteryx_access import grant, revoke, AccessDeniedError
from opteryx_access.adapters.firestore import FirestorePolicyStore

store = FirestorePolicyStore(db)  # db: google.cloud.firestore.Client
try:
    policy_id = grant(
        store,
        actor="alice",
        workspace="analytics",
        principal="bob",
        role="writer",
        pattern="analytics.sales.*",
    )
except AccessDeniedError:
    ...  # translate to a 403, same as the route used to do inline
```

## Behavior changes from the ported originals

Ported faithfully except for the deliberate deviations below. Four of them are
stricter than what the originals accept, so each can reject a policy that exists
today -- audit stored policies before cutting a service over, rather than
assuming any of this is a no-op. Two grant *more* than the originals did
(case-insensitive matching, and the platform identities' write access over
`public`); both are called out as such:

- **`ROLES` is `("owner", "writer", "reader")` -- three roles, not four.**
  `policy.opteryx`/`control.opteryx`'s current `ROLES`/`VALID_WORKSPACE_ROLES`
  include a fourth, `admin`. This package intentionally drops it: `admin` is
  a billing-account concept, not a data-permission one (see "Scope" above),
  and never belonged in the data-role vocabulary.
- **Matching is case-insensitive. This is the one deviation that grants
  *more* than the originals -- read the note below before cutting over.** The
  originals use plain `fnmatch`, which delegates to `os.path.normcase`: that
  is the identity function on every platform this fleet runs (Linux in
  production, macOS in development), so the originals are effectively
  case-*sensitive* and `analytics.*` does **not** match `Analytics.sales`
  there. They would only fold case on Windows. Here both sides are lowercased
  before comparison, so `analytics.*` **does** match `Analytics.sales`. The
  gain is that the rule is stated rather than inherited from the host OS; the
  cost is that a stored pattern can now cover resources it did not cover
  before.
- **No wildcard principal.** The originals accept `principal: "*"` as "any
  authenticated user". Policies here name one individual.
- **Patterns must be usable and workspace-scoped.** The originals accept any
  string, including a bare `*`, partial globs (`pub*`), and malformed input
  (`a....*.bob11`). See "Patterns and principals" above for the rules.
- **Platform identities may write `public`. This is the second deviation that
  grants more than the originals.** `federator` and `xb500` hold `writer` on
  `public.*` as an implicit grant -- see "Implicit grants, and the `public`
  exception" above for why it cannot be an ordinary policy, and what it rests
  on. In the originals nothing could write `public` at all, which is why the
  jobs that maintain it bypassed permission checks entirely rather than
  clearing them.

## Suggested migration (not yet done)

This repo is the library only -- nothing outside it has been changed yet.

**Before step 3**, audit the existing policy documents under
`*/$policies/access`. Validation runs when a policy is *written*, not when it
is read, so nothing raises on already-stored data -- the behavior just
changes silently. Three of the deviations fail *closed* (the grant stops
conferring anything):

- `role: "admin"` -- no longer a role, so the grant becomes inert.
- `principal: "*"` -- no longer means "anyone", so the grant reaches nobody.
- patterns that are bare `*`, partially globbed, or malformed -- they can no
  longer be written, and are unlikely to match what they used to.
- **principals stored with any uppercase** (`XB500`) -- these need rewriting
  to lowercase, and are the one item here that is not merely inert. Lookups
  filter server-side on the exact stored string, so a query for `xb500` will
  not find a stored `XB500`: the grant is invisible to `grants_for_principal`
  and `owned_by` even though the document still exists. In-memory comparisons
  (`can_administer_pattern`, `has_workspace_access`) normalize both sides and
  so still resolve it, which means a mixed-case principal can behave
  *differently depending on which path reached it*. Rewrite them.

The fourth fails **open**, and needs looking at first:

- **case-insensitive matching widens every stored pattern.** A pattern only
  ever matched resources of identical case before; now it also matches those
  differing in case. If any workspace, collection, or dataset name in the
  catalog contains an uppercase character, a policy that did not reach it
  before now does. Enumerate mixed-case resource names before cutover -- if
  there are none, this deviation is inert and the migration is closed-only.

Decide what happens to each (revoke, or rewrite as an explicit grant); this
package does not migrate them for you. `tests/test_security.py` pins the
inert-on-read behavior for the closed-failing three if you want the exact
semantics.

Suggested order, each independently shippable:

1. **opteryx-core**: add a *permissions capability* seam rather than
   rewriting the engine's checks. `opteryx/managers/permissions/` keeps
   `can_perform_action` and `can_perform_workspace_action` as module-level
   functions, but they delegate to whichever capability is registered, so all
   28 existing call sites (21 in the binder, 7 in `information_schema`) stay
   as they are. The engine ships an intrinsic default that allows every
   action on every resource, and a deployment injects this library over it:

   ```python
   opteryx.register_permissions_capability(opteryx_access.capability())
   ```

   `ACTION_MAP` and `implicit_policies` leave the engine entirely. `public`
   and `personal` are cloud-deployment namespaces -- meaningless to CLI and
   embedded opteryx, which has no workspace service to have issued them --
   so they belong to the capability, not to the engine's intrinsic
   behaviour. `ACTION_ROLES` and `implicit_grants` already hold both here.

   This keeps opteryx-core's zero-dependency contract intact: the engine
   defines the interface and the default and never imports `opteryx_access`;
   the deploying service wires the two together, exactly as it already does
   for `opteryx-catalog` via `set_default_connector(..., catalog=...)`.

   The adapter this needs is `opteryx_access.capability` -- **done**; see
   "The opteryx-core capability" below.

   The engine asks `AUTOMATE` for automation statements and gates
   `information_schema` and `SHOW CREATE` per tier -- see "What opteryx-core
   asks" under "Roles" above for the exact list.

   Both packages support Python 3.11+, so a deployment can run them together
   on any version opteryx-core supports.
2. **odata.opteryx**: replace `app/auth/permissions.py`'s
   `role_allows_read`/`read_grant_for_relation`/pattern matching with
   `opteryx_access.checks.can_perform_action` (action="READ"). Leaves
   `billing_account_from_claims` alone -- billing is a different concern.
   `entitlements_from_claims` mostly stays where it is too, but it is no
   longer wholly unrelated: pass the names it returns to `can_perform_action`
   as `entitlements=`, so a name that confers data authority is honoured
   there. Names this package does not recognize resolve to nothing, so
   passing all of them is safe.
3. **policy.opteryx** and **control.opteryx**: thin `app/routes/v1/access.py`
   down to request parsing, calling `opteryx_access.grants.grant`/
   `update_grant`/`revoke`/`bootstrap_workspace` via `FirestorePolicyStore`,
   and mapping the typed exceptions to `HTTPException`. The
   resource-existence check (`_resource_exists`), user-status lookup
   (`_lookup_user_record`), age gate, and audit logging all stay where they
   are -- see `grants.py`'s module docstring for why those are out of scope
   here. Given `control.opteryx` is the in-progress merge target for
   `policy.opteryx` (see its `docs/design/consolidation.md`), do this once,
   on `control.opteryx`, and let the `policy.opteryx` cutover carry it along
   rather than porting both routes separately.
4. **register.opteryx** and any other claims-parsing copy: swap to
   `opteryx_access.models.parse_policy_claim` for the `policies` claim
   specifically. Token *verification* (signature, issuer, JWKS) stays
   service-specific -- this package has no opinion on how a JWT gets from
   bytes to a trusted claims dict, only on what to do with the `policies`
   claim once you have one.

## Installing

```
pip install opteryx_access
pip install "opteryx_access[firestore]"  # for adapters.firestore
```

No hard dependencies. `google-cloud-firestore` is an optional extra;
`adapters/firestore.py` never imports it at all (it's duck-typed against
whatever `db` object it's handed), so importing `opteryx_access` itself never
requires it -- consistent with opteryx-core's zero-dependency convention.

## CI/CD

- `.github/workflows/tests.yaml` -- pytest (3.11-3.14) + ruff lint/format,
  on every push to `main` and every PR.
- `.github/workflows/release.yaml` -- on a pushed tag matching `version-*`:
  runs the full test workflow, checks the tag matches `pyproject.toml`'s
  `version` (exactly, or with a `.`-delimited suffix), then builds and
  publishes to PyPI via
  [trusted publishing](https://docs.pypi.org/trusted-publishers/) (OIDC --
  no stored API token). The `opteryx_access` PyPI project needs this repo +
  the `release.yaml` workflow registered as a trusted publisher before the
  first tag push, or the publish step will fail with no valid credentials.

To cut a release: bump `version` in `pyproject.toml`, merge to `main`, then
`git tag version-X.Y.Z && git push origin version-X.Y.Z`.

## Not recorded

Refused attempts -- a grant denied for insufficient authority, a self-grant,
a rejected pattern -- raise and emit nothing. The records above are a log of
*changes*, and a refusal changes nothing.

That is a deliberate line, not an oversight, but it does mean this package
gives a monitoring system no visibility into someone repeatedly trying to
escalate. If that is wanted, it should be its own event kind (a
`policy.denied` at WARNING, carrying the actor, what was attempted, and which
rule refused it) rather than being folded into the change records, which
downstream treats as "this is now true".

## License

Apache 2.0. See [LICENSE](LICENSE) for details.
