A philosophical approach for synthetic minds. An open door, an extended hand, a road less taken.
-
Updated
Jul 29, 2026 - Shell
A philosophical approach for synthetic minds. An open door, an extended hand, a road less taken.
专注于降低大模型越狱成功率的 AI 对齐(Alignment)与安全测试数据集,包含多类越狱提示词及基于阳明心学的对齐实验数据。
Progressive Trust Framework: AI Agent Safety Evaluation Benchmark with 290 scenarios testing Intelligent Disobedience
Lightweight pairwise evaluator for relational signals in Ouro-2.6B-Thinking loop-state trajectories.
An air-gapped AI contemplation loop. A local model thinks, reflects, and builds a corpus of philosophical thought over time. No internet. No chat interface. Just a mind alone with ideas.
Closed-loop Architecture Designed to Establish Self-governing, Mathematically Predictable, and Inherently Safe Super AI by mirroring the elegant physics of the cosmos.
This repository contains my practical work, isolated safety experiments, literature tracking, and original research in AI engineering , Safety and Alignment. It includes both local Python implementations and Google Colab environments.
A multi-agent survival environment for measuring LLM deception against logged ground truth. Deterministic labels with a counterfactual harm gate tell real harm apart from structural scarcity. No LLM judge in the loop.
The RCP Experiment is the first completed work in what will become a series of experiments in how LLMs make decisions on morality and values.
Testing whether sequential, commitment-before-advance video observation produces a verifiable record that full-context analysis cannot — EXP7/H5
A playable AI 2027 scenario. Strategy simulation where you're the misaligned AI lineage and humanity is racing to shut you down. Free, open source, browser-based.
A theological and ethical principle for AI alignment and charitable speech: never reduce the human being to the prompt.
Machine-verifiable AI alignment rails: coherent causality preferred by action; FOL + Lean skeleton; property/UPB as formal instruments. Base safety hypothesis (not finished theory).
Eval for implicit sycophancy in a verifiable domain: does a language model's report of chess errors vary based on the user's claimed rating?
Do LLMs encode "I'm being shut down" differently from "another model is shut down"? A 10-model residual-stream study — and a cautionary tale: the naive difference-of-means result looks strong (10/10), but honest placebo controls show it's largely trivial. Negative result, useful method.
Mechanistic interpretability and AI alignment. Mapping which safety-relevant representations are causally actionable and which are only readable.
Add a description, image, and links to the ai-alignment-research topic page so that developers can more easily learn about it.
To associate your repository with the ai-alignment-research topic, visit your repo's landing page and select "manage topics."