GetNotes Tools
semantica-agi/semantica
Tool นี้คืออะไร
Semantica เป็นโครงสร้างพื้นฐานแบบ Graph-Native สำหรับระบบ AI ที่ต้องการบริบทและความรับผิดชอบ ช่วยให้องค์กรสามารถนำเข้าข้อมูล สร้างกราฟบริบทและกราฟความรู้ เพื่อการวิเคราะห์และการให้เหตุผลเชิงสาเหตุ พร้อมบันทึกที่มาของการตัดสินใจ ทำให้ AI มีความโปร่งใส ตรวจสอบได้ และน่าเชื่อถือ เหมาะสำหรับทีมที่ทำงานในโดเมนที่มีความเสี่ยงสูงและถูกควบคุม
ข้อมูลโปรเจกต์
ดาว
5.1K
Forks
552
License
MIT
อัปเดต GitHub ล่าสุด
11 ส.ค. 2569
เพิ่มใน GetNotes
17 ส.ค. 2569
Repository
semantica-agi/semantica
เหมาะกับงาน
เหมาะกับอาชีพ
Ecosystem
Python
แปลและเรียบเรียงโดย AI
เนื้อหาฉบับภาษาไทย
ใช้อ่านเพื่อทำความเข้าใจเบื้องต้น โปรดตรวจสอบรายละเอียดสำคัญกับเอกสารต้นฉบับด้านล่าง
โครงสร้างพื้นฐานแบบ Graph-Native สำหรับบริบทและระบบ AI ที่ตรวจสอบได้
Palantir แบบโอเพนซอร์สสำหรับ AI Agents
นำเข้าข้อมูลองค์กรของคุณ, ดึงสิ่งที่สำคัญ, สร้าง Context Graph และ Knowledge Graph (KG), และรันการวิเคราะห์กราฟและการให้เหตุผลเชิงสาเหตุทั้งหมดนี้ พร้อมด้วยการบันทึกที่มาของการตัดสินใจอย่างครบถ้วน ออกแบบมาให้สามารถอธิบายได้, ตรวจสอบย้อนกลับได้ และน่าเชื่อถือ
ระบบอัจฉริยะในการตัดสินใจ · การจัดการบริบท · การให้เหตุผลแบบกำหนดได้ · การจัดการ Ontology · การสร้างแบบจำลองความรู้ · การตรวจสอบย้อนกลับได้ตั้งแต่ต้นจนจบ
โอเพนซอร์ส · ติดตั้งเองได้ · ตรวจสอบได้ · มีการกำกับดูแล · ไม่มีผูกขาดผู้ขาย
การจัดเก็บกราฟแบบ Polyglot · รองรับ RDF และ LPG · มาตรฐาน W3C · ทำงานร่วมกันได้
สร้างขึ้นสำหรับโดเมนที่มีความเสี่ยงสูงและถูกควบคุม
pip install semanticaKnowledge Explorer · Context Graphs · Reasoning Engine · Decision Intelligence · Ontology Hub
AI agents ส่วนใหญ่ทำงานโดยไม่มีร่องรอย พวกมันจัดเก็บ embeddings ไม่ใช่ความหมาย: บริบทที่ไม่สามารถอธิบายได้, การตัดสินใจที่ไม่สามารถตรวจสอบได้ ในธุรกิจสินเชื่อ ช่องว่างนี้คือความเสี่ยงด้านการปฏิบัติตามข้อกำหนด ไม่ใช่แค่ความไม่สะดวก: การอนุมัติของ underwriting agent จะต้องสามารถตอบคำถาม "ทำไม" ของหน่วยงานกำกับดูแลได้ในอีกหลายเดือนข้างหน้า
Semantica ทำหน้าที่เป็นชั้นโครงสร้างพื้นฐานแบบกำหนดได้ที่อยู่ใต้ LLM, vector store และ agent framework ของคุณ: ไม่จำเป็นต้องใช้ LLM สำหรับการสร้างกราฟ, การให้เหตุผล หรือการบันทึกที่มา
⚠️ ความสามารถในการอธิบายระดับระบบ ไม่ใช่ความสามารถในการอธิบายระดับโมเดลพื้นฐาน Semantica ไม่ได้เปิดเผยหรือสร้างสิ่งที่เกิดขึ้น ภายใน LLM ขึ้นมาใหม่ — การให้เหตุผลภายในหรือ chain-of-thought ของมันยังคงไม่โปร่งใส เช่นเดียวกับระบบภายนอกใดๆ Semantica อธิบายสิ่งที่อยู่ ภายนอก โมเดล: บริบทและข้อมูลที่ป้อนเข้าไป, การตัดสินใจที่เกิดขึ้น, ที่มาของมัน, ความสัมพันธ์ที่เกี่ยวข้อง, นโยบายที่ใช้ และร่องรอยการดำเนินการทั้งหมด
เหมาะสำหรับใคร:
- ทีมแพลตฟอร์ม AI/ML ที่พัฒนา agents ซึ่งทำการตัดสินใจที่มีผลกระทบสำคัญ และต้องการบริบทที่มีโครงสร้าง สามารถสอบถามได้ ซึ่งสร้างขึ้นจากข้อมูลดิบที่กระจัดกระจาย ไม่ใช่แค่ดัชนีเวกเตอร์
- ทีมแพลตฟอร์มข้อมูลบน Databricks หรือ Snowflake ที่ต้องการเปลี่ยนตารางที่มีอยู่ใน Unity Catalog หรือ Snowflake warehouse ให้เป็น knowledge graph ที่มีการกำกับดูแลและติดตามสายข้อมูล โดยไม่ต้องส่งออกข้อมูลนั้นไปยัง SaaS ของบุคคลที่สามก่อน
- ทีมงานด้านการปฏิบัติตามข้อกำหนด, ความเสี่ยง และการตรวจสอบ ที่ต้องการคำตอบที่ชัดเจนสำหรับคำถาม "ทำไม AI ถึงทำเช่นนั้น?" ในรูปแบบที่หน่วยงานกำกับดูแลจะยอมรับได้จริง
- องค์กรที่อยู่ภายใต้การกำกับดูแล (การเงิน, การดูแลสุขภาพ, กฎหมาย, รัฐบาล, การป้องกันประเทศ) ที่ไม่สามารถส่งมอบระบบแบบ black box และไม่สามารถส่งข้อมูลของตนไปยัง SaaS ของผู้อื่นเพื่อใช้งานได้
- วิศวกรแพลตฟอร์มและโครงสร้างพื้นฐาน ที่ต้องการให้ KG, การให้เหตุผล และ provenance stack สามารถติดตั้งเองได้และสลับเปลี่ยนได้ ไม่ถูกผูกติดกับแบ็กเอนด์ของผู้ขายรายเดียว
- วิศวกรข้อมูลและความรู้ ที่สร้าง KG จากข้อมูลที่ยุ่งเหยิงและมาจากหลายแหล่ง: เอนทิตีและความสัมพันธ์จะถูกดึงออกมา, ข้อเท็จจริงที่ขัดแย้งกันจะถูกทำเครื่องหมายแทนที่จะถูกเขียนทับอย่างเงียบๆ, และข้อมูลซ้ำซ้อนจะถูกรวมเข้าด้วยกันก่อนที่จะกลายเป็นข้อมูลรบกวน
เริ่มต้นใช้งานด่วน · สถาปัตยกรรม · สิ่งที่คุณจะได้รับ · ทำไมต้อง Semantica · ระบบอัจฉริยะในการตัดสินใจ · Context Graphs · สูตร: Audit Trail · การอ้างอิงโมดูล · การผสานรวม · CLI · ประสิทธิภาพ · การติดตั้ง
สิ่งที่คุณจะได้รับจาก Semantica
- Context Graphs: กราฟที่มีโครงสร้างและสามารถสอบถามได้ของทุกสิ่งที่ agent ของคุณรู้, ตัดสินใจ และให้เหตุผล
- Decision Intelligence: ทุกการตัดสินใจเป็นอ็อบเจกต์ระดับเฟิร์สคลาส: ตรวจสอบย้อนกลับได้, ค้นหาได้จากกรณีตัวอย่าง และเชื่อมโยงเชิงสาเหตุ
- AI Governance & Ontology: ข้อจำกัด SHACL, การตรวจจับความขัดแย้ง, กฎการปฏิบัติตามข้อกำหนด, การสร้าง OWL และการจัดการคำศัพท์ SKOS พร้อมด้วยตัวแก้ไขแบบภาพ
- ความสามารถในการตรวจสอบอย่างเต็มรูปแบบ: W3C PROV-O provenance บนทุกข้อเท็จจริง พร้อม audit trails ที่สามารถส่งออกเป็น JSON, CSV หรือ RDF
- Deterministic Reasoning: Forward chaining, Rete network, Datalog และ SPARQL พร้อมเส้นทางที่อธิบายได้อย่างสมบูรณ์ ไม่ใช่ black boxes
- Knowledge Pipeline: การนำเข้าจากหลายแหล่ง, การแบ่งส่วนข้อมูลที่รับรู้เอนทิตี, การดึง NER/ความสัมพันธ์/เหตุการณ์ และการสร้าง knowledge graph พร้อมการ
Semantica เสริมการทำงานของสแต็กที่คุณมีอยู่แล้ว แทนที่จะเข้ามาแทนที่ คุณสามารถเก็บ LLM, vector store และ agent framework ของคุณไว้ได้เหมือนเดิม Semantica จะเพิ่มบันทึกการตัดสินใจ, การให้เหตุผลเชิงสาเหตุ, ที่มาของข้อมูล, การกำกับดูแลออนโทโลยี, การตรวจจับความขัดแย้ง และบันทึกการตรวจสอบเพิ่มเติมเข้ามา เอ็นจิ้นการให้เหตุผล, การสร้าง KG และเลเยอร์ที่มาของข้อมูลนั้นเป็นแบบกำหนดผลลัพธ์ได้ทั้งหมด (deterministic) โดยไม่จำเป็นต้องใช้ LLM ในการใช้งาน
เริ่มต้นใช้งานอย่างรวดเร็ว
pip install semanticafrom semantica.context import ContextGraph
graph = ContextGraph(advanced_analytics=True)
# การตัดสินใจทุกครั้งของเอเจนต์จะกลายเป็นโหนดความรู้ที่สามารถสอบถามและตรวจสอบได้
decision_id = graph.record_decision(
category="vendor_selection",
scenario="Choose cloud provider for HIPAA workload",
reasoning="AWS offers BAA, mature HIPAA tooling, and existing team expertise",
outcome="selected_aws",
confidence=0.93,
)
# ถามว่า "ทำไมสิ่งนี้ถึงเกิดขึ้น?" และรับคำตอบที่เป็นโครงสร้างจริง
chain = graph.trace_decision_chain(decision_id) # ลำดับบรรพบุรุษเชิงสาเหตุทั้งหมด
similar = graph.find_similar_decisions("cloud vendor", max_results=5) # กรณีตัวอย่าง
impact = graph.analyze_decision_impact(decision_id) # แผนที่ผลกระทบปลายน้ำ
compliant = graph.check_decision_rules({"category": "vendor_selection"}) # เกตตรวจสอบนโยบายตรวจสอบการติดตั้งของคุณใน 5 วินาที:
semantica doctor
# Python 3.11.9 pass
# semantica 0.6.5 pass
# faiss vector store pass
# Config file pass ~/.semantica/config.yamlหาก Semantica ช่วยแก้ปัญหาจริงให้คุณ การกดดาวจะช่วยให้ผู้อื่นค้นพบได้
⭐ กดดาวบน GitHub · เข้าร่วม Discord
สถาปัตยกรรม
Semantica เป็นไปป์ไลน์แบบ end-to-end ที่แท้จริง ไม่ใช่แค่ไลบรารีเดียวที่มีชื่อทางการตลาด ทุกขั้นตอนด้านล่างเป็นโมดูลที่พร้อมใช้งานและสามารถนำเข้าได้อย่างอิสระ:
Sources → Ingest → Parse → Normalize → Split → Extract → Conflict Detection → Deduplication
→ Knowledge Graph → [ Ontology · Reasoning · Provenance · Decisions ] → Enriched KG
→ Vector Store + Polyglot Graph Store (RDF & LPG) → Export / Visualize / REST · MCP · CLI- Ingest: ไฟล์, เว็บ, ฐานข้อมูล, แพลตฟอร์มข้อมูลองค์กร (Databricks, Snowflake), คลาวด์ (Google Drive, Elasticsearch), สตรีม (Kafka, Kinesis), Git, อีเมล, MCP
- Parse → Normalize → Split: การแยกวิเคราะห์เอกสาร, การทำให้ข้อความ/เอนทิตี/วันที่เป็นมาตรฐาน, การแบ่งส่วนข้อมูลที่รับรู้เอนทิตีแบบ GraphRAG-native
- Extract → Conflict Detection → Deduplication: NER, ความสัมพันธ์, เหตุการณ์, ทริปเปิล; ข้อเท็จจริงที่ขัดแย้งกันจะถูกตั้งค่าสถานะและแก้ไขก่อนที่จะรวมเข้าด้วยกัน
- Knowledge Graph:
GraphBuilderสร้างกราฟ; ข้อเท็จจริงแบบ bi-temporal และการวิเคราะห์กราฟเต็มรูปแบบ (centrality, communities, link prediction) ทำงานอยู่บนกราฟนี้ - Ontology · Reasoning · Provenance · Decisions: เลเยอร์อัจฉริยะที่อยู่บน KG พร้อมการกำกับดูแล SHACL/OWL, การอนุมาน Rete/Datalog/SPARQL, ลำดับวงศ์ตระกูล W3C PROV-O และบันทึกการตัดสินใจระดับเฟิร์สคลาส
- Storage: ออกแบบมาให้เป็น polyglot โดยมี RDF triple stores (Oxigraph แบบฝัง, Blazegraph, Apache Jena, Eclipse RDF4J), Labeled Property Graphs (Neo4j, FalkorDB, Apache AGE, AWS Neptune) และ vector stores ซึ่งทั้งหมดสามารถสลับเปลี่ยนได้โดยไม่ต้องแก้ไขโค้ดของคุณ
- Outputs: การส่งออก (RDF, OWL, Parquet, Cypher, JSON-LD), การแสดงภาพแบบโต้ตอบ และการเข้าถึงผ่าน REST API, MCP server หรือ CLI
→ แผนภาพ Mermaid แบบเต็มสำหรับไปป์ไลน์และวงจรชีวิตของ Decision Intelligence
Decision Intelligence
Decision Intelligence เปลี่ยนทุกการตัดสินใจของ AI จากการอนุมานชั่วคราวให้เป็นบันทึกถาวรที่สามารถตรวจสอบและสอบถามได้ มันตอบคำถามที่ว่า "AI ของคุณตัดสินใจอะไร ทำไม และเกิดอะไรขึ้นต่อไป?": คำถามที่หน่วยงานกำกับดูแลและทีมบริหารความเสี่ยงขององค์กรถามด้วยความเร่งด่วนที่เพิ่มขึ้น
ใน Semantica การตัดสินใจไม่ใช่แค่บรรทัดบันทึก (log line) แต่เป็นโหนดกราฟระดับเฟิร์สคลาสที่มีวงจรชีวิตที่สมบูรณ์ ในโดเมนที่มีการควบคุม ทุกการตัดสินใจของ AI จะต้องสามารถตรวจสอบย้อนกลับไปยังแหล่งที่มาและสามารถป้องกันได้ต่อผู้ตรวจสอบ: record_decision() สร้างบันทึกถาวรที่มีโครงสร้างที่สามารถส่งออกเป็น W3C PROV-O ซึ่งเป็นรูปแบบที่กรอบการทำงานด้านการปฏิบัติตามข้อกำหนดส่วนใหญ่ยอมรับสำหรับการส่งให้หน่วยงานกำกับดูแล
record_decision() → จัดเก็บเป็นโหนดกราฟพร้อมบริบทที่มีโครงสร้างครบถ้วน
add_causal_relationship() → เชื่อมโยงกับสาเหตุต้นน้ำและผลกระทบปลายน้ำ
find_similar_decisions() → การค้นหาตัวอย่างเชิงความหมายจากการตัดสินใจที่ผ่านมาทั้งหมด
trace_decision_chain() → ลำดับบรรพบุรุษเชิงสาเหตุทั้งหมดกลับไปยังสาเหตุหลัก
analyze_decision_impact() → แผนที่ผลกระทบปลายน้ำ - ทุกสิ่งที่การตัดสินใจนี้ส่งผลกระทบ
check_decision_rules() → เกตตรวจสอบการปฏิบัติตามนโยบายเทียบกับชุดกฎที่กำหนดค่าได้
export / audit trail → W3C PROV-O, CSV หรือ JSON สำหรับการส่งให้หน่วยงานกำกับดูแลfrom semantica.context import ContextGraph
graph = ContextGraph(advanced_analytics=True)
# บันทึกการตัดสินใจพร้อมบริบทที่มีโครงสร้างครบถ้วน
app_id = graph.record_decision(
category="credit_application",
scenario="Personal loan, $85k income, 31% DTI, 3yr employment",
reasoning="Income meets threshold; employment stable; no adverse credit events",
outcome="proceed_to_underwriting",
confidence=0.88,
metadata={"applicant_id": "A-7291"},
)
uw_id = graph.record_decision(
category="loan_underwriting",
scenario="Underwriting review for A-7291",
reasoning="DTI within policy; clean 36-month credit history",
outcome="approved",
confidence=0.94,
)
rate_id = graph.record_decision(
category="interest_rate",
scenario="Rate assignment for approved loan A-7291",
outcome="rate_set_8.9pct",
reasoning="Prime + 2.4% based on risk tier B2",
confidence=0.99,
)
# สร้างห่วงโซ่สาเหตุที่ตรวจสอบได้ - relationship_type ต้องเป็นหนึ่งใน
# CAUSED, INFLUENCED, หรือ PRECEDENT_FOR
graph.add_causal_relationship(app_id, uw_id, relationship_type="CAUSED")
graph.add_causal_relationship(uw_id, rate_id, relationship_type="INFLUENCED")
# สอบถามข้อมูลอัจฉริยะ
chain = graph.trace_decision_chain(rate_id)
similar = graph.find_similar_decisions("personal loan approval, 31% DTI", max_results=5)
impact = graph.analyze_decision_impact(uw_id)
compliant = graph.check_decision_rules({"category": "loan_underwriting", "confidence": 0.94})
insights = graph.get_decision_insights()Context Graphs
Context Graph คือเลเยอร์หน่วยความจำที่มีโครงสร้างที่ RAG แบบดั้งเดิมขาดหายไป แทนที่จะเป็น embeddings แบบแบนที่ตอบคำถามว่า "อะไรที่คล้ายกัน?" Context Graph จะตอบคำถามว่า "อะไรที่เชื่อมโยงกัน ทำไม และอย่างไร?" ทุกเอนทิตี ความสัมพันธ์ การตัดสินใจ และข้อเท็จจริงเป็นโหนดระดับเฟิร์สคลาสที่สามารถสอบถามได้ด้วยการสำรวจกราฟ เอนทิตีเชื่อมโยงกับเอกสารต้นฉบับ การตัดสินใจเชื่อมโยงกับหลักฐานและผลที่ตามมา ข้อเท็จจริงมีที่มาของข้อมูลครบถ้วน และความขัดแย้งจะถูกตรวจจับ ไม่ใช่ถูกเขียนทับอย่างเงียบๆ
from semantica.context import ContextGraph, AgentContext
from semantica.vector_store import VectorStore
graph = ContextGraph(advanced_analytics=True)
# เพิ่มโหนดพร้อมคุณสมบัติแบบมีประเภท
graph.add_node("acme_corp", "Organization", name="Acme Corp", industry="SaaS")
graph.add_node("alice_chen", "Person", name="Alice Chen", role="CTO")
graph.add_node("contract_001", "Contract", value=2_400_000, currency="USD")
# เพิ่มขอบแบบมีประเภทและมีน้ำหนัก (kwargs เพิ่มเติมจะกลายเป็นเมตาดาต้าของขอบ)
graph.add_edge("alice_chen", "acme_corp", edge_type="works_for", since="2019-03-01")
graph.add_edge("acme_corp", "contract_001", edge_type="party_to", signed="2024-01-15")
# การสำรวจแบบ BFS - กระโดดผ่านกราฟจากโหนดใดก็ได้
neighbors = graph.get_neighbors("acme_corp", hops=2)
# สแนปช็อต ณ จุดเวลา - กราฟในสภาพที่เป็นอยู่ ณ วันที่ในอดีตใดๆ
snapshot = graph.state_at("2024-01-01")
# AgentContext - API ระดับสูงสำหรับเวิร์กโฟลว์หน่วยความจำของเอเจนต์
vs = VectorStore(backend="faiss")
ctx = AgentContext(vector_store=vs, knowledge_graph=graph)
ctx.store("Alice approved the Acme renewal in Q1 2024", conversation_id="conv_001")
retrieved = ctx.retrieve("who approved the Acme contract?")ทำไมต้องใช้กราฟแทน embeddings: การสำรวจกราฟพบการเชื่อมโยงที่ embeddings พลาดไป (บุคคลที่อยู่ห่างจากสัญญา 3 hops); ทุกโหนดมีที่มาของข้อมูลเพื่อให้คุณสามารถถามได้เสมอว่า "สิ่งนี้มาจากไหน?"; ความขัดแย้งจะถูกตั้งค่าสถานะก่อนที่จะทำให้ฐานความรู้ของคุณเสียหาย; สแนปช็อต ณ จุดเวลาช่วยให้คุณสามารถเล่นประวัติซ้ำได้โดยไม่ต้องประมวลผลใหม่
สูตร: บันทึกการตรวจสอบสำหรับการตัดสินใจที่มีการควบคุม
รูปแบบหลัก: บันทึกห่วงโซ่การตัดสินใจที่เชื่อมโยงกันเชิงสาเหตุ แนบที่มาของข้อมูลกับทุกเอนทิตี และส่งออกบันทึกการตรวจสอบที่พร้อมสำหรับหน่วยงานกำกับดูแล
from semantica.context import ContextGraph
from semantica.provenance import ProvenanceManager
from semantica.export import RDFExporter
graph = ContextGraph(advanced_analytics=True)
prov = ProvenanceManager(storage_path="./audit.db")
# บันทึกห่วงโซ่การตัดสินใจ
d1 = graph.record_decision(
category="drug_interaction_check", scenario="Patient P-4821: warfarin + amiodarone co-prescribed",
reasoning="Amiodarone potentiates warfarin's anticoagulant effect", outcome="flag_for_review", confidence=0.91,
)
d2 = graph.record_decision(
category="dosage_adjustment", scenario="INR monitoring plan for P-4821",
reasoning="Reduce warfarin dose per interaction severity; recheck INR in 5 days", outcome="dose_reduced_30pct", confidence=0.87,
)
# relationship_type ต้องเป็นหนึ่งใน CAUSED, INFLUENCED, หรือ PRECEDENT_FOR
graph.add_causal_relationship(d1, d2, relationship_type="CAUSED")
# ติดตามที่มาของข้อมูลสำหรับทุกเอนทิตี
prov.track_entity("patient_P4821", source="ehr/medication_orders_2024.json",
metadata={"extractor": "NamedEntityRecognizer"})
# ส่งออก W3C PROV-O สำหรับการส่งให้หน่วยงานกำกับดูแล - RDFExporter คาดหวัง
# {"entities": [...], "relationships": [...]}, ดังนั้นจึงต้องแมป ContextGraph.to_dict()'s
# {"nodes": [...], "edges": [...]} ให้เป็นรูปแบบนี้ก่อน
graph_dict = graph.to_dict()
kg = {
"entities": [{"id": n["id"], "type": n["type"], "text": n["content"]} for n in graph_dict["nodes"]],
"relationships": [
{"source_id": e["source"], "target_id": e["target"], "type": e["type"]}
for e in graph_dict["edges"]
],
}
RDFExporter().export(kg, "audit_trail.ttl", format="turtle")
ส่วนที่ 3/4
สูตรเพิ่มเติม (ไปป์ไลน์ GraphRAG, เอ็นจิ้นกฎ AML, การแปลง Ontology เป็น KG ในครั้งเดียว) อยู่ใน **[สูตรเพิ่มเติม](#more-recipes)** ด้านล่าง
---
## สำรวจแพลตฟอร์ม
แต่ละโมดูลด้านล่างสามารถนำเข้าได้โดยอิสระ พร้อมตัวอย่างโค้ดที่ใช้งานได้จริงซึ่งได้รับการตรวจสอบกับโครงสร้างซอร์สปัจจุบัน คุณสามารถใช้หนึ่งหรือทั้งหมดก็ได้
| โมดูล | สิ่งที่ทำ |
| --- | --- |
| [`semantica.ingest`](#semanticaingest-multi-source-ingestion) | ไฟล์, เว็บ, ฐานข้อมูล, API, สตรีม, อีเมล, Git, Parquet, Databricks, Snowflake, MCP |
| [`semantica.semantic_extract`](#semanticasemantic_extract-ner-relations-events-triplets) | NER, การสกัดความสัมพันธ์, การตรวจจับเหตุการณ์, การสร้างทริปเปิล |
| [`semantica.kg`](#semanticakg-knowledge-graph-construction--analysis) | การสร้างกราฟ, Centrality, ชุมชน, การทำนายลิงก์ |
| [`semantica.reasoning`](#semanticareasoning-forward-chaining-rete-datalog-sparql) | Forward chaining, Rete, Datalog, SPARQL, อธิบายได้ทั้งหมด |
| [`semantica.vector_store`](#semanticavector_store-hybrid--filtered-semantic-search) | FAISS, Qdrant, Weaviate, Milvus, Pinecone, PgVector, การค้นหาแบบไฮบริด |
| [`semantica.split`](#semanticasplit-graphrag-native-document-chunking) | การแบ่งส่วนเอกสารที่รับรู้เอนทิตี, ความสัมพันธ์, ออนโทโลยี สำหรับ GraphRAG |
| [`semantica.provenance`](#semanticaprovenance-w3c-prov-o-lineage) | W3C PROV-O lineage บนทุกข้อเท็จจริง |
| [`semantica.ontology`](#semanticaontology-owl-generation-shacl-validation) | การสร้าง OWL, การตรวจสอบ SHACL, คำศัพท์ SKOS |
| [`semantica.conflicts`](#semanticaconflicts-conflict-detection--resolution) | ตรวจจับและแก้ไขข้อเท็จจริงที่ขัดแย้งกันจากหลายแหล่ง |
| [`semantica.deduplication`](#semanticadeduplication-entity-resolution-at-scale) | การแก้ไขเอนทิตีในขนาดใหญ่ |
| [`semantica.normalize`](#semanticanormalize-data-normalization--cleaning) | การทำให้ข้อความ, เอนทิตี, วันที่, และตัวเลขเป็นมาตรฐาน; การทำความสะอาดชุดข้อมูล |
| [`semantica.pipeline`](#semanticapipeline-pipeline-dsl) | DSL ไปป์ไลน์แบบประกาศ, ขนาน สำหรับ ingest → extract → build → export |
| [`semantica.export`](#semanticaexport-rdf-owl-parquet-cypher-json-ld) | RDF, OWL, Parquet, Cypher, JSON-LD |
| [`semantica.visualization`](#semanticavisualization-interactive-graph-workbench) | กราฟแบบแรงดึงดูด, ลำดับชั้นออนโทโลยี, แดชบอร์ดเชิงเวลา |
| [Temporal Intelligence](#temporal-intelligence-bi-temporal-graphs--time-travel) | ข้อเท็จจริงแบบสองเวลา, Allen interval algebra, การเดินทางข้ามเวลา |
| [Multi-Agent (Agno)](#multi-agent-shared-context-with-agno) | กราฟบริบทที่ใช้ร่วมกันหนึ่งเดียวสำหรับทุกเอเจนต์ในทีม |
**↓ ขยาย [การอ้างอิงโมดูล](#module-reference)** ด้านล่างสำหรับตัวอย่างการทำงานของทุกโมดูล หรือข้ามไปที่ [สูตรเพิ่มเติม](#more-recipes), เมทริกซ์ [การผสานรวม](#integrations) แบบเต็ม, [รายการเครื่องมือ MCP](#mcp-server), และ [REST endpoints](#rest-api)
---
## การอ้างอิงโมดูล
ขยายโมดูลใดๆ ด้านล่างเพื่อดูตัวอย่างที่สามารถรันได้
<details>
<summary><b><code>semantica.ingest</code></b>: การนำเข้าจากหลายแหล่ง</summary>
<a id="semanticaingest-multi-source-ingestion"></a>
นำเข้าจากไฟล์, เว็บ, ฐานข้อมูล, API, สตรีม, อีเมล, Git repos, Parquet, Databricks, Snowflake, หรือ MCP servers ทั้งหมดผ่านอินเทอร์เฟซแบบรวมศูนย์
```python
from semantica.ingest import FileIngestor, WebIngestor, ParquetIngestor, DBIngestor
# นำเข้าไดเรกทอรีทั้งหมดของสัญญา (PDF, DOCX, HTML, TXT)
docs = FileIngestor().ingest_directory("./contracts/", recursive=True)
# นำเข้าเนื้อหาเว็บสดพร้อมการปฏิบัติตาม robots.txt
pages = WebIngestor().ingest_url("https://example.com/reports/annual-2024.html")
# นำเข้าข้อมูลที่มีโครงสร้างจาก Parquet พร้อมการบีบอัด Snappy
records = ParquetIngestor().ingest("./data/transactions.parquet")
# นำเข้าจากฐานข้อมูล SQL - ระบุตารางที่จะดึง
rows = DBIngestor().ingest_database(
connection_string="postgresql://user:pass@localhost/mydb",
include_tables=["customer_events"],
max_rows_per_table=50_000,
)# แพลตฟอร์มข้อมูลระดับองค์กร - ดึงตารางโดยตรงจาก lakehouse
# หรือ warehouse ของคุณ พร้อม lineage แทนที่จะส่งออกเป็น CSV ก่อน
from semantica.ingest import DatabricksIngestor, SnowflakeIngestor
# pip install "semantica[db-databricks]"
databricks = DatabricksIngestor(
host="https://adb-xxx.azuredatabricks.net",
token="dapi-xxxxxxxx", # หรือ client_id/client_secret สำหรับ OAuth M2M
http_path="/sql/1.0/warehouses/xxxxxxxx",
catalog="main",
)
customers = databricks.ingest_table("customers", limit=10_000)
sales = databricks.ingest_query("SELECT * FROM sales WHERE region = 'EMEA'")
table_lineage = databricks.get_table_lineage("customers", catalog="main", schema="default") # Unity Catalog lineage
# pip install semantica[db-snowflake]
snowflake = SnowflakeIngestor(
account="myaccount",
user="myuser",
password="mypassword", # หรือ private_key=... สำหรับ key-pair; ใช้ authenticator="oauth", token=... สำหรับ OAuth
warehouse="COMPUTE_WH",
database="MYDB",
)
orders = snowflake.ingest_table("ORDERS", limit=10_000)หมายเหตุความปลอดภัย: ห้ามฮาร์ดโค้ดข้อมูลรับรอง (token, password, private_key) ในโค้ดที่ใช้งานจริง; ให้ส่งผ่านตัวแปรสภาพแวดล้อม (เช่น DATABRICKS_TOKEN, SNOWFLAKE_PASSWORD) หรือตัวจัดการความลับ
แหล่งที่มารองรับ: ไฟล์ในเครื่อง (PDF, DOCX, PPTX, HTML, TXT, CSV, JSON, YAML, Excel, XML) · หน้าเว็บ · ฟีด RSS/Atom · REST APIs · ฐานข้อมูล (PostgreSQL, MySQL, SQLite, Oracle, SQL Server) · ชุดข้อมูล Parquet · Databricks (Unity Catalog + Delta Lake) · Snowflake · Git repositories · อีเมล (IMAP/POP3) · สตรีมข้อความ (Kafka, RabbitMQ, Kinesis, Pulsar) · ทรัพยากร MCP · Apache Arrow/Feather/IPC (ArrowIngestor)
DuckDB, Elasticsearch, Google Drive, HuggingFace, MongoDB, และ Pandas ingestion ก็มีให้ใช้งาน (DuckDBIngestor, ElasticIngestor, GDriveIngestor, HuggingFaceIngestor, MongoIngestor, PandasIngestor) แต่ยังไม่ได้ถูก re-export จาก semantica.ingest namespace ระดับบนสุด — ให้ import โดยตรง: from semantica.ingest.duckdb_ingestor import DuckDBIngestor
สกัดความรู้ที่มีโครงสร้างจากข้อความดิบในครั้งเดียว
from semantica.semantic_extract import (
NamedEntityRecognizer,
RelationExtractor,
EventDetector,
TripletExtractor,
)
text = """
Anthropic CEO Dario Amodei announced a $7.3B Series E funding round in partnership
with Google and Spark Capital, valuing the company at $61.5B as of Q4 2024.
"""
# การรู้จำเอนทิตีที่มีชื่อพร้อมการกำหนดเกณฑ์ความเชื่อมั่น
ner = NamedEntityRecognizer(confidence_threshold=0.7)
entities = ner.extract_entities(text)
# → [Entity(name="Dario Amodei", type="PERSON"), Entity(name="Anthropic", type="ORG"),
# Entity(name="Google", type="ORG"), Entity(name="$7.3B", type="MONEY"), ...]
# การสกัดความสัมพันธ์ - รองรับสองทิศทาง
rel_extractor = RelationExtractor(confidence_threshold=0.6, bidirectional=True)
relations = rel_extractor.extract_relations(text, entities=entities)
# → [Relation(subject="Dario Amodei", predicate="ceo_of", object="Anthropic"),
# Relation(subject="Anthropic", predicate="raised", object="$7.3B Series E"), ...]
# การตรวจจับเหตุการณ์พร้อมการประมวลผลเชิงเวลา
events = EventDetector(extract_participants=True, extract_time=True).detect_events(text)
# → [Event(type="FUNDING", participants=["Anthropic","Google","Spark Capital"],
# amount="$7.3B", date="Q4 2024")]
# RDF triplets พร้อมเมตาดาต้า provenance เสริม
triplets = TripletExtractor(include_temporal=True, include_provenance=True).extract_triplets(text)
# → [("Anthropic", "valuation", "$61.5B"), ("Dario Amodei", "is_ceo_of", "Anthropic"), ...]การประมวลผลแบบแบตช์ในเอกสารจำนวนมากใช้ ner.process_batch([...]) ไม่ใช่ extract_entities_batch ต่อการเรียกบนคลาส facade
สร้าง Knowledge Graph สำหรับการผลิตจากเอกสารและรันอัลกอริทึมกราฟบนนั้น
from semantica.ingest import FileIngestor
from semantica.kg import (
GraphBuilder,
GraphAnalyzer,
CentralityCalculator,
CommunityDetector,
PathFinder,
LinkPredictor,
BiTemporalFact,
)
from datetime import datetime
# สร้าง KG - รวมเอนทิตีที่ซ้ำกัน, ติดตามขอบเชิงเวลา
sources = FileIngestor().ingest_directory("./contracts/", recursive=True)
kg = GraphBuilder(merge_entities=True, enable_temporal=True).build(sources)
# การวิเคราะห์กราฟ
analyzer = GraphAnalyzer()
analysis = analyzer.analyze_graph(kg) # เมตริกกราฟทั้งหมด
centrality = CentralityCalculator()
degree = centrality.calculate_degree_centrality(kg) # เอนทิตีที่มีการเชื่อมต่อมากที่สุด
betweenness = centrality.calculate_betweenness_centrality(kg)
communities = CommunityDetector().detect_communities(kg, method="louvain") # กลุ่มธรรมชาติ
path = PathFinder().find_shortest_path(kg, "alice_chen", "contract_001")
predictions = LinkPredictor().predict_links(kg, top_k=10) # การทำนายความสัมพันธ์
# ข้อเท็จจริงแบบสองเวลา - ติดตามเวลาที่ถูกต้องเทียบกับเวลาที่บันทึกไว้แยกกัน
fact = BiTemporalFact(
valid_from=datetime(2024, 3, 1),
valid_until=datetime(2025, 1, 1),
recorded_at=datetime(2024, 3, 5),
)
```</details>
</details>
<details>
<summary><b><code>semantica.reasoning</code></b>: การอนุมานแบบ Forward Chaining, Rete, Datalog, SPARQL</summary>
<a id="semanticareasoning-forward-chaining-rete-datalog-sparql"></a>
รันการอนุมานแบบใช้กฎที่อธิบายได้ ไม่ใช่กล่องดำ
```python
from semantica.reasoning import ReteEngine, Rule, Fact, RuleType
rete = ReteEngine()
rete.build_network([
Rule(
rule_id="aml_flag",
name="Flag high-risk transactions",
conditions=[
{"field": "amount", "operator": ">", "value": 10_000},
{"field": "country", "operator": "in", "value": ["IR", "KP", "SY"]},
],
conclusion="flag_for_compliance_review",
rule_type=RuleType.IMPLICATION,
),
Rule(
rule_id="velocity_check",
name="Flag rapid sequential transfers",
conditions=[
{"field": "transfers_in_1h", "operator": ">", "value": 5},
{"field": "total_amount", "operator": ">", "value": 50_000},
],
conclusion="flag_velocity_breach",
rule_type=RuleType.IMPLICATION,
),
])
rete.add_fact(Fact("tx_001", "transaction", [{"amount": 15_000, "country": "IR"}]))
flagged = rete.match_patterns()
# → [{"rule": "aml_flag", "matched_facts": ["tx_001"], "conclusion": "flag_for_compliance_review"}]ข้อจำกัดปัจจุบัน: ตัวจับคู่เงื่อนไข alpha-node ของ ReteEngine ถูกออกแบบให้เรียบง่ายโดยเจตนาในเวอร์ชันนี้ — โปรดตรวจสอบเอาต์พุตของ match_patterns() เทียบกับชุดกฎจริงของคุณก่อนที่จะนำไปใช้กับเกตการปฏิบัติตามข้อกำหนดในการผลิต; การประเมินเงื่อนไขที่เลือกสรรมากขึ้นอยู่ในแผนงาน
# Datalog แบบเรียกซ้ำ - ภาษามนุษย์สำหรับคิวรีกราฟ
from semantica.reasoning import DatalogReasoner
engine = DatalogReasoner()
engine.add_fact("parent(tom, bob)")
engine.add_fact("parent(bob, ann)")
engine.add_fact("parent(ann, pat)")
engine.add_rule("ancestor(X, Y) :- parent(X, Y).")
engine.add_rule("ancestor(X, Z) :- parent(X, Y), ancestor(Y, Z).")
ancestors = engine.query("ancestor(tom, ?X)")
# → [{"X": "bob"}, {"X": "ann"}, {"X": "pat"}]# การอนุมานที่อธิบายได้ - ติดตามเส้นทาง ไม่ใช่แค่คำตอบ
from semantica.reasoning import ExplanationGenerator, Reasoner
reasoner = Reasoner()
reasoner.add_fact("parent(tom, bob)")
reasoner.add_rule("ancestor(X, Y) :- parent(X, Y)")
result = reasoner.forward_chain()
explainer = ExplanationGenerator()
explanation = explainer.generate_explanation(result)
# → Explanation(conclusion="...", steps=[ReasoningStep(...)], justification=Justification(...))Vector store ที่ใช้งานง่ายพร้อมแบ็กเอนด์หลายตัว, การค้นหาแบบไฮบริด และการดึงข้อมูลที่คำนึงถึงการตัดสินใจ
from semantica.vector_store import VectorStore, HybridSearch
# แสดงแบ็กเอนด์แบบ In-memory ที่นี่: HybridSearch และ explain_decision() ใช้งานได้ทันทีเอกสารโปรเจกต์
อ่านเอกสารต้นฉบับ
README วิธีติดตั้ง วิธีใช้งาน และข้อกำหนดจาก repository ต้นฉบับ

Graph-Native Infrastructure for Context and Accountable AI Systems
The Open Source Palantir for AI Agents
Ingest your enterprise data, extract what matters, build a Context Graph and knowledge graph (KG), and run graph analytics and causal reasoning over all of it, with full decision provenance baked in. Explainable, traceable, and trustworthy by design.
Decision Intelligence · Context Management · Deterministic Reasoning · Ontology Management · Knowledge Modeling · End-to-End Traceability
Open Source · Self-Hostable · Auditable · Governed · Zero Vendor Lock-In
Polyglot Graph Storage · RDF & LPG Support · W3C Standards · Interoperable
Built for High-Stakes, Regulated Domains
pip install semantica
Knowledge Explorer · Context Graphs · Reasoning Engine · Decision Intelligence · Ontology Hub
▶ Watch the full platform walkthrough
Most AI agents act without a trail. They store embeddings, not meaning: context that can't be explained, decisions that can't be audited. In lending, that gap is a compliance exposure, not an inconvenience: an underwriting agent's approval has to survive a regulator's "why" months later.
Semantica sits underneath your LLM, vector store, and agent framework as a deterministic infrastructure layer: no LLM required for graph construction, reasoning, or provenance.
⚠️ System-level explainability, not foundation-model explainability. Semantica does not expose or reconstruct what happens inside the LLM — its internal reasoning or chain-of-thought stays opaque, as it does for any external system. Semantica explains what's outside the model: the context and data fed in, the decision produced, its provenance, relevant relationships, applied policies, and the full execution trail.
Who it's for:
- AI/ML platform teams shipping agents that make consequential decisions and need structured, queryable context built from fragmented raw data, not just a vector index
- Data platform teams on Databricks or Snowflake who need to turn tables already sitting in Unity Catalog or a Snowflake warehouse into a governed, lineage-tracked knowledge graph, without exporting that data to a third-party SaaS first
- Compliance, risk, and audit teams who need a straight answer to "why did the AI do that?" in a format a regulator will actually accept
- Regulated enterprises (finance, healthcare, legal, government, defense) that can't ship a black box, and can't send their data to someone else's SaaS to get one
- Platform and infra engineers who want the KG, reasoning, and provenance stack self-hosted and swappable, not locked to one vendor's backend
- Data and knowledge engineers building a KG from messy, multi-source data: entities and relationships get extracted, conflicting or contradictory facts are flagged instead of silently overwritten, and duplicates are merged before they turn into noise
Quick Start · Architecture · What You Get · Why Semantica · Decision Intelligence · Context Graphs · Recipe: Audit Trail · Module Reference · Integrations · CLI · Performance · Install
What Semantica Gives You
- Context Graphs: A structured, queryable graph of everything your agent knows, decides, and reasons about
- Decision Intelligence: Every decision is a first-class object: traceable, searchable by precedent, and causally linked
- AI Governance & Ontology: SHACL constraints, conflict detection, compliance rules, OWL generation, and SKOS vocabulary management with a visual editor
- Full Auditability: W3C PROV-O provenance on every fact, with audit trails exportable to JSON, CSV, or RDF
- Deterministic Reasoning: Forward chaining, Rete network, Datalog, and SPARQL with fully explainable paths, not black boxes
- Knowledge Pipeline: Multi-source ingestion, entity-aware chunking, NER/relation/event extraction, and knowledge graph construction, with semantic deduplication and provenance-preserving merges throughout
- Enterprise Data Platforms: Native connectors for Databricks (Unity Catalog + Delta Lake, PAT/OAuth M2M auth, catalog/schema/table/lineage introspection) and Snowflake (warehouse/database/schema, key-pair and OAuth auth), so tables already living in your lakehouse or warehouse become graph nodes with provenance, not another export/import hop
- Graph Analytics: Centrality, community detection, link prediction, and shortest-path queries over the graph you just built
- Polyglot Graph Storage: Native RDF (embedded Oxigraph, Blazegraph, Apache Jena, Eclipse RDF4J via SPARQL) and Labeled Property Graphs (Neo4j, FalkorDB, Apache AGE, AWS Neptune via Cypher), plus vector stores, all swappable without touching your code
- Visualization: Explore any graph, ontology, or timeline in an interactive browser workbench
- Drop-in Integrations: Native Agno and CrewAI support, a full-featured MCP server, a comprehensive CLI, a REST API, and plugins across major editors
Why Semantica
| Vector DB + RAG | Plain LLM Memory | Semantica | |
|---|---|---|---|
| Recall method | Embedding similarity | Token window | Graph traversal + semantic search |
| Decision history | Not stored | Not stored | First-class queryable objects |
| Provenance | None | None | W3C PROV-O, source-linked |
| Reasoning | None | Black box | Forward chain, Rete, Datalog, SPARQL |
| Conflict detection | Silent overwrite | Silent overwrite | Detected, flagged, resolved |
| Time travel | No | No | Point-in-time graph snapshots |
| Compliance export | None | None | PROV-O, SHACL, OWL, RDF |
| Policy enforcement | None | None | Built-in rule engine + SHACL |
| Entity resolution | No | No | Blocking + semantic deduplication |
| Multi-agent context | Separate per agent | Separate per agent | Single shared intelligence layer |
Semantica complements your existing stack rather than replacing it. Keep your LLM, vector store, and agent framework exactly as they are; Semantica adds the decision records, causal reasoning, provenance, ontology governance, conflict detection, and audit trails on top. The reasoning engines, KG construction, and provenance layer are fully deterministic; no LLM is required to use them.
Quick Start
pip install semanticafrom semantica.context import ContextGraph
graph = ContextGraph(advanced_analytics=True)
# Every agent decision becomes a queryable, auditable knowledge node
decision_id = graph.record_decision(
category="vendor_selection",
scenario="Choose cloud provider for HIPAA workload",
reasoning="AWS offers BAA, mature HIPAA tooling, and existing team expertise",
outcome="selected_aws",
confidence=0.93,
)
# Ask "why did this happen?" and get a real, structured answer
chain = graph.trace_decision_chain(decision_id) # full causal ancestry
similar = graph.find_similar_decisions("cloud vendor", max_results=5) # precedents
impact = graph.analyze_decision_impact(decision_id) # downstream influence map
compliant = graph.check_decision_rules({"category": "vendor_selection"}) # policy gateVerify your install in 5 seconds:
semantica doctor
# Python 3.11.9 pass
# semantica 0.6.5 pass
# faiss vector store pass
# Config file pass ~/.semantica/config.yamlIf Semantica solves a real problem for you, a star helps others find it.
⭐ Star on GitHub · Join Discord
Architecture
Semantica is a real end-to-end pipeline, not a single library with a marketing name. Every stage below is a shipping module, independently importable:
Sources → Ingest → Parse → Normalize → Split → Extract → Conflict Detection → Deduplication
→ Knowledge Graph → [ Ontology · Reasoning · Provenance · Decisions ] → Enriched KG
→ Vector Store + Polyglot Graph Store (RDF & LPG) → Export / Visualize / REST · MCP · CLI- Ingest: files, web, databases, enterprise data platforms (Databricks, Snowflake), cloud (Google Drive, Elasticsearch), streams (Kafka, Kinesis), Git, email, MCP
- Parse → Normalize → Split: document parsing, text/entity/date normalization, GraphRAG-native entity-aware chunking
- Extract → Conflict Detection → Deduplication: NER, relations, events, triplets; conflicting facts flagged and resolved before they merge
- Knowledge Graph:
GraphBuilderconstructs the graph; bi-temporal facts and full graph analytics (centrality, communities, link prediction) run on top of it - Ontology · Reasoning · Provenance · Decisions: the intelligence layer sitting on the KG, with SHACL/OWL governance, Rete/Datalog/SPARQL inference, W3C PROV-O lineage, and first-class decision records
- Storage: polyglot by design, with RDF triple stores (embedded Oxigraph, Blazegraph, Apache Jena, Eclipse RDF4J), Labeled Property Graphs (Neo4j, FalkorDB, Apache AGE, AWS Neptune), and vector stores, all swappable without touching your code
- Outputs: export (RDF, OWL, Parquet, Cypher, JSON-LD), interactive visualization, and access via REST API, MCP server, or CLI
→ Full Mermaid diagrams for the pipeline and the decision intelligence lifecycle
Decision Intelligence
Decision Intelligence turns every AI choice from an ephemeral inference into a permanent, auditable, queryable record. It answers "what did your AI decide, why, and what happened next?": the question regulators and enterprise risk teams ask with increasing urgency.
In Semantica, a decision is not a log line. It is a first-class graph node with a full lifecycle. In regulated domains, every AI decision must be traceable to a source and defensible to an auditor: record_decision() creates a permanent, structured record exportable as W3C PROV-O, the format most compliance frameworks accept for regulator submission.
record_decision() → stored as a graph node with full structured context
add_causal_relationship() → linked to upstream causes and downstream effects
find_similar_decisions() → semantic precedent search across all past decisions
trace_decision_chain() → full causal ancestry back to root causes
analyze_decision_impact() → downstream influence map - everything this decision affected
check_decision_rules() → policy compliance gate against configurable rule sets
export / audit trail → W3C PROV-O, CSV, or JSON for regulator submissionfrom semantica.context import ContextGraph
graph = ContextGraph(advanced_analytics=True)
# Record decisions with full structured context
app_id = graph.record_decision(
category="credit_application",
scenario="Personal loan, $85k income, 31% DTI, 3yr employment",
reasoning="Income meets threshold; employment stable; no adverse credit events",
outcome="proceed_to_underwriting",
confidence=0.88,
metadata={"applicant_id": "A-7291"},
)
uw_id = graph.record_decision(
category="loan_underwriting",
scenario="Underwriting review for A-7291",
reasoning="DTI within policy; clean 36-month credit history",
outcome="approved",
confidence=0.94,
)
rate_id = graph.record_decision(
category="interest_rate",
scenario="Rate assignment for approved loan A-7291",
outcome="rate_set_8.9pct",
reasoning="Prime + 2.4% based on risk tier B2",
confidence=0.99,
)
# Build the auditable causal chain - relationship_type must be one of
# CAUSED, INFLUENCED, or PRECEDENT_FOR
graph.add_causal_relationship(app_id, uw_id, relationship_type="CAUSED")
graph.add_causal_relationship(uw_id, rate_id, relationship_type="INFLUENCED")
# Query the intelligence
chain = graph.trace_decision_chain(rate_id)
similar = graph.find_similar_decisions("personal loan approval, 31% DTI", max_results=5)
impact = graph.analyze_decision_impact(uw_id)
compliant = graph.check_decision_rules({"category": "loan_underwriting", "confidence": 0.94})
insights = graph.get_decision_insights()Context Graphs
A Context Graph is the structured memory layer that traditional RAG is missing. Instead of flat embeddings that answer "what is similar?", a Context Graph answers "what is connected, why, and how?" Every entity, relationship, decision, and fact is a first-class node, queryable by graph traversal. Entities link to source documents, decisions link to evidence and consequences, facts carry full provenance, and conflicts are detected, not silently overwritten.
from semantica.context import ContextGraph, AgentContext
from semantica.vector_store import VectorStore
graph = ContextGraph(advanced_analytics=True)
# Add nodes with typed properties
graph.add_node("acme_corp", "Organization", name="Acme Corp", industry="SaaS")
graph.add_node("alice_chen", "Person", name="Alice Chen", role="CTO")
graph.add_node("contract_001", "Contract", value=2_400_000, currency="USD")
# Add typed, weighted edges (extra kwargs become edge metadata)
graph.add_edge("alice_chen", "acme_corp", edge_type="works_for", since="2019-03-01")
graph.add_edge("acme_corp", "contract_001", edge_type="party_to", signed="2024-01-15")
# BFS traversal - hop through the graph from any node
neighbors = graph.get_neighbors("acme_corp", hops=2)
# Point-in-time snapshot - the graph as it existed on any past date
snapshot = graph.state_at("2024-01-01")
# AgentContext - high-level API for agent memory workflows
vs = VectorStore(backend="faiss")
ctx = AgentContext(vector_store=vs, knowledge_graph=graph)
ctx.store("Alice approved the Acme renewal in Q1 2024", conversation_id="conv_001")
retrieved = ctx.retrieve("who approved the Acme contract?")Why graph over embeddings: traversal finds connections embeddings miss (a person 3 hops from a contract); every node carries provenance so you can always ask "where did this come from?"; conflicts are flagged before they corrupt your knowledge base; point-in-time snapshots let you replay history without reprocessing.
Recipe: Audit Trail for a Regulated Decision
The flagship pattern: record a causally-linked decision chain, attach provenance to every entity, and export a regulator-ready audit trail.
from semantica.context import ContextGraph
from semantica.provenance import ProvenanceManager
from semantica.export import RDFExporter
graph = ContextGraph(advanced_analytics=True)
prov = ProvenanceManager(storage_path="./audit.db")
# Record the decision chain
d1 = graph.record_decision(
category="drug_interaction_check", scenario="Patient P-4821: warfarin + amiodarone co-prescribed",
reasoning="Amiodarone potentiates warfarin's anticoagulant effect", outcome="flag_for_review", confidence=0.91,
)
d2 = graph.record_decision(
category="dosage_adjustment", scenario="INR monitoring plan for P-4821",
reasoning="Reduce warfarin dose per interaction severity; recheck INR in 5 days", outcome="dose_reduced_30pct", confidence=0.87,
)
# relationship_type must be one of CAUSED, INFLUENCED, or PRECEDENT_FOR
graph.add_causal_relationship(d1, d2, relationship_type="CAUSED")
# Track provenance for every entity
prov.track_entity("patient_P4821", source="ehr/medication_orders_2024.json",
metadata={"extractor": "NamedEntityRecognizer"})
# Export W3C PROV-O for regulator submission - to_kg_dict() is the official
# adapter that emits the {"entities": [...], "relationships": [...]} /
# source_id shape RDFExporter expects, so no manual field mapping is needed
kg = graph.to_kg_dict()
RDFExporter().export(kg, "audit_trail.ttl", format="turtle")More recipes (GraphRAG pipelines, an AML rules engine, ontology-to-KG in one pass) are in More Recipes below.
Explore the Platform
Every module below is independently importable, with working code samples verified against the current source tree; use one or all of them.
| Module | What it does |
|---|---|
semantica.ingest | Files, web, databases, APIs, streams, email, Git, Parquet, Databricks, Snowflake, MCP |
semantica.semantic_extract | NER, relation extraction, event detection, triplet generation |
semantica.kg | Graph construction, centrality, communities, link prediction |
semantica.reasoning | Forward chaining, Rete, Datalog, SPARQL, fully explainable |
semantica.vector_store | FAISS, Qdrant, Weaviate, Milvus, Pinecone, PgVector, hybrid search |
semantica.split | Entity-aware, relation-aware, ontology-aware chunking for GraphRAG |
semantica.provenance | W3C PROV-O lineage on every fact |
semantica.ontology | OWL generation, SHACL validation, SKOS vocabularies |
semantica.conflicts | Detect and resolve conflicting facts across sources |
semantica.deduplication | Entity resolution at scale |
semantica.normalize | Text, entity, date, and number normalization; dataset cleaning |
semantica.pipeline | Declarative, parallel pipeline DSL for ingest → extract → build → export |
semantica.export | RDF, OWL, Parquet, Cypher, JSON-LD |
semantica.visualization | Force-directed graphs, ontology hierarchies, temporal dashboards |
| Temporal Intelligence | Bi-temporal facts, Allen interval algebra, time travel |
| Multi-Agent (Agno) | One shared context graph across every agent on a team |
↓ Expand Module Reference below for every module's working example, or jump to More Recipes, the full Integrations matrix, MCP tool list, and REST endpoints.
Module Reference
Expand any module below for its runnable example.
Ingest from files, web, databases, APIs, streams, email, Git repos, Parquet, Databricks, Snowflake, or MCP servers, all through a unified interface.
from semantica.ingest import FileIngestor, WebIngestor, ParquetIngestor, DBIngestor
# Ingest an entire directory of contracts (PDF, DOCX, HTML, TXT)
docs = FileIngestor().ingest_directory("./contracts/", recursive=True)
# Ingest live web content with robots.txt compliance
pages = WebIngestor().ingest_url("https://example.com/reports/annual-2024.html")
# Ingest structured data from Parquet with Snappy compression
records = ParquetIngestor().ingest("./data/transactions.parquet")
# Ingest from a SQL database - specify which tables to pull
rows = DBIngestor().ingest_database(
connection_string="postgresql://user:pass@localhost/mydb",
include_tables=["customer_events"],
max_rows_per_table=50_000,
)# Enterprise data platforms - pull tables straight out of your lakehouse
# or warehouse, with lineage, instead of exporting to CSV first
from semantica.ingest import DatabricksIngestor, SnowflakeIngestor
# pip install "semantica[db-databricks]"
databricks = DatabricksIngestor(
host="https://adb-xxx.azuredatabricks.net",
token="dapi-xxxxxxxx", # or client_id/client_secret for OAuth M2M
http_path="/sql/1.0/warehouses/xxxxxxxx",
catalog="main",
)
customers = databricks.ingest_table("customers", limit=10_000)
sales = databricks.ingest_query("SELECT * FROM sales WHERE region = 'EMEA'")
table_lineage = databricks.get_table_lineage("customers", catalog="main", schema="default") # Unity Catalog lineage
# pip install semantica[db-snowflake]
snowflake = SnowflakeIngestor(
account="myaccount",
user="myuser",
password="mypassword", # or private_key=... for key-pair; use authenticator="oauth", token=... for OAuth
warehouse="COMPUTE_WH",
database="MYDB",
)
orders = snowflake.ingest_table("ORDERS", limit=10_000)Security Note: Never hardcode credentials (token, password, private_key) in production code; pass them via environment variables (e.g., DATABRICKS_TOKEN, SNOWFLAKE_PASSWORD) or a secrets manager.
Supported sources: Local files (PDF, DOCX, PPTX, HTML, TXT, CSV, JSON, YAML, Excel, XML) · Web pages · RSS/Atom feeds · REST APIs · Databases (PostgreSQL, MySQL, SQLite, Oracle, SQL Server) · Parquet datasets · Databricks (Unity Catalog + Delta Lake) · Snowflake · Git repositories · Email (IMAP/POP3) · Message streams (Kafka, RabbitMQ, Kinesis, Pulsar) · MCP resources · Apache Arrow/Feather/IPC (ArrowIngestor)
DuckDB, Elasticsearch, Google Drive, HuggingFace, MongoDB, and Pandas ingestion also ship (DuckDBIngestor, ElasticIngestor, GDriveIngestor, HuggingFaceIngestor, MongoIngestor, PandasIngestor) but aren't re-exported from the top-level semantica.ingest namespace yet — import them directly: from semantica.ingest.duckdb_ingestor import DuckDBIngestor.
Extract structured knowledge from raw text in one pass.
from semantica.semantic_extract import (
NamedEntityRecognizer,
RelationExtractor,
EventDetector,
TripletExtractor,
)
text = """
Anthropic CEO Dario Amodei announced a $7.3B Series E funding round in partnership
with Google and Spark Capital, valuing the company at $61.5B as of Q4 2024.
"""
# Named entity recognition with confidence thresholding
ner = NamedEntityRecognizer(confidence_threshold=0.7)
entities = ner.extract_entities(text)
# → [Entity(name="Dario Amodei", type="PERSON"), Entity(name="Anthropic", type="ORG"),
# Entity(name="Google", type="ORG"), Entity(name="$7.3B", type="MONEY"), ...]
# Relationship extraction - bidirectional support
rel_extractor = RelationExtractor(confidence_threshold=0.6, bidirectional=True)
relations = rel_extractor.extract_relations(text, entities=entities)
# → [Relation(subject="Dario Amodei", predicate="ceo_of", object="Anthropic"),
# Relation(subject="Anthropic", predicate="raised", object="$7.3B Series E"), ...]
# Event detection with temporal processing
events = EventDetector(extract_participants=True, extract_time=True).detect_events(text)
# → [Event(type="FUNDING", participants=["Anthropic","Google","Spark Capital"],
# amount="$7.3B", date="Q4 2024")]
# RDF triplets with optional provenance metadata
triplets = TripletExtractor(include_temporal=True, include_provenance=True).extract_triplets(text)
# → [("Anthropic", "valuation", "$61.5B"), ("Dario Amodei", "is_ceo_of", "Anthropic"), ...]Batch processing across many documents uses ner.process_batch([...]), not a per-call extract_entities_batch on the facade class.
Build a production knowledge graph from documents and run graph algorithms over it.
from semantica.ingest import FileIngestor
from semantica.kg import (
GraphBuilder,
GraphAnalyzer,
CentralityCalculator,
CommunityDetector,
PathFinder,
LinkPredictor,
BiTemporalFact,
)
from datetime import datetime
# Build KG - merge duplicate entities, track temporal edges
sources = FileIngestor().ingest_directory("./contracts/", recursive=True)
kg = GraphBuilder(merge_entities=True, enable_temporal=True).build(sources)
# Graph analytics
analyzer = GraphAnalyzer()
analysis = analyzer.analyze_graph(kg) # full graph metrics
centrality = CentralityCalculator()
degree = centrality.calculate_degree_centrality(kg) # most-connected entities
betweenness = centrality.calculate_betweenness_centrality(kg)
communities = CommunityDetector().detect_communities(kg, method="louvain") # natural clusters
path = PathFinder().find_shortest_path(kg, "alice_chen", "contract_001")
predictions = LinkPredictor().predict_links(kg, top_k=10) # relationship predictions
# Bi-temporal facts - track valid time vs. recorded time independently
fact = BiTemporalFact(
valid_from=datetime(2024, 3, 1),
valid_until=datetime(2025, 1, 1),
recorded_at=datetime(2024, 3, 5),
)Run explainable rule-based inference, not a black box.
from semantica.reasoning import ReteEngine, Rule, Fact, RuleType
rete = ReteEngine()
rete.build_network([
Rule(
rule_id="aml_flag",
name="Flag high-risk transactions",
conditions=[
{"field": "amount", "operator": ">", "value": 10_000},
{"field": "country", "operator": "in", "value": ["IR", "KP", "SY"]},
],
conclusion="flag_for_compliance_review",
rule_type=RuleType.IMPLICATION,
),
Rule(
rule_id="velocity_check",
name="Flag rapid sequential transfers",
conditions=[
{"field": "transfers_in_1h", "operator": ">", "value": 5},
{"field": "total_amount", "operator": ">", "value": 50_000},
],
conclusion="flag_velocity_breach",
rule_type=RuleType.IMPLICATION,
),
])
rete.add_fact(Fact("tx_001", "transaction", [{"amount": 15_000, "country": "IR"}]))
flagged = rete.match_patterns()
# → [{"rule": "aml_flag", "matched_facts": ["tx_001"], "conclusion": "flag_for_compliance_review"}]Current limitation: ReteEngine's alpha-node condition matcher is intentionally simple in this release — validate match_patterns() output against your actual rule set before wiring it into a production compliance gate; more selective condition evaluation is on the roadmap.
# Recursive Datalog - natural language for graph queries
from semantica.reasoning import DatalogReasoner
engine = DatalogReasoner()
engine.add_fact("parent(tom, bob)")
engine.add_fact("parent(bob, ann)")
engine.add_fact("parent(ann, pat)")
engine.add_rule("ancestor(X, Y) :- parent(X, Y).")
engine.add_rule("ancestor(X, Z) :- parent(X, Y), ancestor(Y, Z).")
ancestors = engine.query("ancestor(tom, ?X)")
# → [{"X": "bob"}, {"X": "ann"}, {"X": "pat"}]# Explainable reasoning - trace the path, not just the answer
from semantica.reasoning import ExplanationGenerator, Reasoner
reasoner = Reasoner()
reasoner.add_fact("parent(tom, bob)")
reasoner.add_rule("ancestor(X, Y) :- parent(X, Y)")
result = reasoner.forward_chain()
explainer = ExplanationGenerator()
explanation = explainer.generate_explanation(result)
# → Explanation(conclusion="...", steps=[ReasoningStep(...)], justification=Justification(...))Drop-in vector store with multiple backends, hybrid search, and decision-aware retrieval.
from semantica.vector_store import VectorStore, HybridSearch
# In-memory backend shown here: HybridSearch and explain_decision() work out of the box.
# Swap backend="qdrant" / "weaviate" / "milvus" / "pinecone" / "pgvector" / "faiss" once you
# scale past a single process — search() and store_decision() work identically on all of them.
vs = VectorStore(backend="inmemory", dimension=1536)
# Store a decision with scenario description and outcome
vs.store_decision(
scenario="Personal loan A-7291, $85k income, 31% DTI, 3yr employment",
outcome="approved",
confidence=0.94,
category="loan_underwriting",
)
# Semantic similarity search
results = vs.search(
query="personal loan approval with low DTI",
limit=10,
)
# Hybrid search - dense + sparse retrieval in one pass with RRF fusion
hs = HybridSearch(vector_store=vs)
hits = hs.search("high-risk transactions 2024")
# Explain why a decision was retrieved
explanation = vs.explain_decision(results[0]["id"])Backends: faiss · qdrant · weaviate · milvus · pinecone · pgvector · sqlite · inmemory
KG-aware splitting that preserves entity boundaries, relation triplets, and ontology concepts, essential for GraphRAG pipelines.
from semantica.split import TextSplitter, EntityAwareChunker, RelationAwareChunker
text = open("contracts/master_agreement.txt").read()
# Standard recursive chunking
chunks = TextSplitter(method="recursive", chunk_size=1000, chunk_overlap=200).split(text)
# Entity-aware chunking - never splits a named entity across chunks (GraphRAG)
chunks = TextSplitter(method="entity_aware", ner_method="llm", chunk_size=1000).split(text)
# Relation-aware chunking - preserves (subject, predicate, object) triplets intact
chunks = RelationAwareChunker(chunk_size=1000, preserve_triplets=True).chunk(text)
# Graph-based chunking - uses centrality to find natural community boundaries
chunks = TextSplitter(method="graph_based", chunk_size=1000).split(text)
# Hierarchical chunking - multi-level (section → paragraph → sentence)
chunks = TextSplitter(method="hierarchical", levels=["section", "paragraph"]).split(text)Supported methods: recursive · token · sentence · paragraph · semantic_transformer · entity_aware · relation_aware · graph_based · ontology_aware · hierarchical · community_detection · centrality_based · llm
Every fact is linked to its source. No black boxes, no mystery outputs.
from semantica.provenance import ProvenanceManager
prov = ProvenanceManager(storage_path="./provenance.db")
# Track where every entity came from
prov.track_entity(
entity_id="acme_corp",
source="contracts/acme_master_agreement_2024.pdf",
metadata={"page": 1, "confidence": 0.97, "extractor": "NamedEntityRecognizer"},
)
# Track a relationship's provenance - entity linkage travels in metadata
prov.track_relationship(
relationship_id="alice_works_for_acme",
source="hr_records/employees_q1_2024.csv",
metadata={"source_entity_id": "alice_chen", "target_entity_id": "acme_corp"},
)
# Answer "where did this come from?"
lineage = prov.get_lineage("acme_corp")
trail = prov.trace_lineage("alice_chen") # full ancestor chain
entry = prov.get_provenance("acme_corp")Generate ontologies from data, validate shapes, and manage your vocabulary.
from semantica.ontology import OntologyGenerator, OntologyValidator
data = {
"entities": [
{"id": "acme_corp", "type": "Organization", "industry": "SaaS", "founded": 2012},
{"id": "alice_chen", "type": "Person", "role": "CTO", "since": 2019},
],
"relationships": [
{"source": "alice_chen", "target": "acme_corp", "type": "works_for"},
],
}
gen = OntologyGenerator(base_uri="https://semantica.dev/ontology/")
ontology = gen.generate_ontology(data)
classes = gen.infer_classes(data)
props = gen.infer_properties(data, classes)
optimized = gen.optimize_ontology(ontology)
# Validate against SHACL shapes
validator = OntologyValidator()
report = validator.validate(ontology)
# → ValidationResult(valid=True, consistent=True, satisfiable=True, errors=[], warnings=[])Detect and resolve conflicting facts from multiple sources before they corrupt your knowledge base.
from semantica.conflicts import ConflictDetector, ConflictResolver, SourceTracker
entities_from_source_a = [
{"id": "alice_chen", "role": "CTO", "salary": 250_000, "start_date": "2019-03-01"},
]
entities_from_source_b = [
{"id": "alice_chen", "role": "VP Eng", "salary": 275_000, "start_date": "2019-03-01"},
]
# Detect all conflict types: value, type, relationship, temporal, logical
detector = ConflictDetector()
conflicts = detector.detect_conflicts(entities_from_source_a + entities_from_source_b)
# → [Conflict(entity="alice_chen", field="role", values=["CTO","VP Eng"], severity="HIGH"),
# Conflict(entity="alice_chen", field="salary", values=[250000,275000], severity="MEDIUM")]
# Resolve using multiple strategies
resolver = ConflictResolver()
resolved = resolver.resolve_conflicts(conflicts, strategy="credibility_weighted") # weighted by source trust
resolved = resolver.resolve_conflicts(conflicts, strategy="most_recent") # prefer most recent
resolved = resolver.resolve_conflicts(conflicts, strategy="voting") # majority wins
# Track source credibility over time
tracker = SourceTracker()
tracker.register_source("source_a", source_type="document", credibility_score=0.85)
tracker.register_source("source_b", source_type="document", credibility_score=0.72)Block, cluster, and merge duplicates with semantic similarity.
from semantica.deduplication import DuplicateDetector, EntityMerger
entities = [
{"id": "e1", "name": "Acme Corporation", "domain": "acme.com"},
{"id": "e2", "name": "Acme Corp.", "domain": "acme.com"},
{"id": "e3", "name": "ACME Corp", "domain": "acme.co"},
{"id": "e4", "name": "Globex Industries", "domain": "globex.com"},
]
detector = DuplicateDetector(similarity_threshold=0.75, use_clustering=True)
candidates = detector.detect_duplicates(entities)
groups = detector.detect_duplicate_groups(entities)
# → DuplicateGroup(entities=["e1","e2","e3"], confidence=0.91, strategy="semantic+blocking")
merger = EntityMerger(preserve_provenance=True)
ops = merger.merge_duplicates(entities, strategy="keep_most_complete")
history = merger.get_merge_history()Standardize text, entities, dates, numbers, and encodings before building your knowledge graph.
from semantica.normalize import (
TextNormalizer,
EntityNormalizer,
DateNormalizer,
NumberNormalizer,
DataCleaner,
)
# Unicode, whitespace, casing, HTML tags, smart quotes
text = TextNormalizer().normalize(" Acme Corp.'s Q4 report... ")
# → "Acme Corp.'s Q4 report..."
# Alias resolution + entity disambiguation with confidence scores
canonical = EntityNormalizer().normalize_entity("ACME Corp.")
# → NormalizedEntity(canonical="Acme Corporation", type="Organization", confidence=0.91)
# Natural language date parsing with timezone conversion
dt = DateNormalizer().normalize_date("3 weeks ago")
# → datetime(2026, 7, 1, tzinfo=UTC)
# Unit conversion and currency normalization
price = NumberNormalizer().normalize_number("$1.25M USD")
# → NormalizedNumber(value=1_250_000, currency="USD")
# Deduplicate, validate, and impute missing values across a dataset
clean = DataCleaner().clean_data(records, remove_duplicates=True, handle_missing=True)Compose ingestion, extraction, and graph-building into a declarative, parallel pipeline.
from semantica.pipeline import PipelineBuilder, ExecutionEngine
builder = PipelineBuilder()
# add_step() returns the created PipelineStep, not the builder, so these don't chain
builder.add_step("ingest", step_type="ingest", source="./contracts/", recursive=True)
builder.add_step("extract", step_type="ner_extract")
builder.add_step("relations", step_type="relation_extract")
builder.add_step("build_kg", step_type="kg_build", merge_entities=True)
builder.add_step("deduplicate", step_type="deduplicate", threshold=0.75)
builder.add_step("export", step_type="export", format="turtle", output="kg.ttl")
# connect_steps() and set_parallelism() return the builder, so these do chain
pipeline = (
builder
.connect_steps("ingest", "extract")
.connect_steps("extract", "relations")
.connect_steps("relations", "build_kg")
.connect_steps("build_kg", "deduplicate")
.connect_steps("deduplicate", "export")
.set_parallelism(4)
.build(name="contracts_pipeline")
)
engine = ExecutionEngine()
result = engine.execute_pipeline(pipeline)
status = engine.get_pipeline_status(pipeline.name)
progress = engine.get_progress(pipeline.name)Track when facts were true in the world vs. when they were recorded, and query either axis.
from semantica.context import ContextGraph
from semantica.kg import (
BiTemporalFact,
TemporalGraphQuery,
TemporalNormalizer,
)
from datetime import datetime
graph = ContextGraph(advanced_analytics=True)
graph.add_node("alice_chen", "Person", role="VP Engineering")
graph.add_node("acme_corp", "Organization", valuation=1_200_000_000)
# A temporally-bounded edge - valid_from/valid_until define when it held true
graph.add_edge(
"alice_chen", "acme_corp", edge_type="works_for",
valid_from="2024-03-01T00:00:00", valid_until="2025-01-01T00:00:00",
)
# Point-in-time snapshots - replay history without reprocessing
snapshot_2023 = graph.state_at("2023-06-01")
snapshot_2024 = graph.state_at("2024-01-01")
# Bi-temporal facts - valid_time is when true in the world;
# recorded_at is when you learned about it
fact = BiTemporalFact(
valid_from=datetime(2024, 3, 1),
valid_until=datetime(2025, 1, 1),
recorded_at=datetime(2024, 3, 5),
)
# Query facts valid within a time window - to_kg_dict() is the official
# adapter that emits {"entities", "relationships"} with source_id/target_id
# keys, the shape query_time_range() expects (no manual mapping required)
kg = graph.to_kg_dict()
tq = TemporalGraphQuery()
facts_in_window = tq.query_time_range(
kg, query="valid_facts", start_time="2024-01-01", end_time="2024-12-31"
)
# Normalize natural language temporal expressions - returns a (start, end) range
norm = TemporalNormalizer()
start, end = norm.normalize("last quarter")Export to any format required by regulators, graph databases, or downstream systems.
from semantica.export import (
RDFExporter,
JSONExporter,
ParquetExporter,
LPGExporter,
ReportGenerator,
)
kg = {"entities": [...], "relationships": [...]}
rdf = RDFExporter()
turtle_str = rdf.export_to_rdf(kg, format="turtle") # returns string
jsonld_str = rdf.export_to_rdf(kg, format="json-ld")
rdf.export(kg, "kg_audit.ttl", format="turtle")
rdf.export(kg, "kg_audit.jsonld", format="json-ld")
rdf.export(kg, "kg_audit.nt", format="n-triples")
# Columnar analytics - Snappy-compressed Parquet (writes kg_snapshot_entities.parquet
# and kg_snapshot_relationships.parquet)
ParquetExporter(compression="snappy").export_knowledge_graph(kg, "kg_snapshot")
# JSON knowledge graph
JSONExporter().export_knowledge_graph(kg, "kg.json")
# Neo4j / Memgraph Cypher statements for graph database import
LPGExporter().export(kg, "kg_import.cypher")
# Human-readable HTML report
ReportGenerator().generate_report(
{"title": "KG Audit Report", "summary": "Weekly ingestion summary", "metrics": {"entities": len(kg["entities"])}},
file_path="audit_report.html",
format="html",
)Render force-directed graphs, community maps, ontology hierarchies, and temporal dashboards.
from semantica.visualization import (
KGVisualizer,
OntologyVisualizer,
EmbeddingVisualizer,
TemporalVisualizer,
)
import numpy as np
kg = {"entities": [...], "relationships": [...]}
# Interactive force-directed graph (opens in browser)
viz = KGVisualizer(layout="force", color_scheme="default")
viz.visualize_network(kg, output="interactive", file_path="kg.html")
viz.visualize_communities(kg, communities, output="interactive")
viz.visualize_centrality(kg, centrality, centrality_type="degree")
viz.visualize_entity_types(kg, output="html", file_path="entity_types.html")
# Ontology class hierarchy
OntologyVisualizer().visualize_hierarchy(ontology, output="interactive")
# 2D embedding projection (UMAP / t-SNE / PCA)
EmbeddingVisualizer().visualize_2d_projection(
embeddings=np.array([...]),
labels=["entity_a", "entity_b"],
method="umap",
)
# Timeline scrubber - watch the graph evolve
TemporalVisualizer().visualize_timeline(kg, output="interactive")One shared intelligence layer. All agents read and write to the same context graph.
# pip install semantica[agno]
from agno.agent import Agent
from agno.team import Team
from agno.models.anthropic import Claude
from semantica.context import ContextGraph
from semantica.vector_store import VectorStore
from integrations.agno import AgnoSharedContext, AgnoDecisionKit, AgnoKGToolkit
shared = AgnoSharedContext(
vector_store=VectorStore(backend="faiss"),
knowledge_graph=ContextGraph(advanced_analytics=True),
decision_tracking=True,
)
researcher = Agent(
name="Researcher",
model=Claude(id="claude-sonnet-4-5"),
memory=shared.bind_agent("researcher"),
tools=[AgnoKGToolkit(context=shared)],
)
analyst = Agent(
name="Analyst",
model=Claude(id="claude-sonnet-4-5"),
memory=shared.bind_agent("analyst"),
tools=[AgnoDecisionKit(context=shared)],
)
team = Team(agents=[researcher, analyst], mode="coordinate")
# Researcher's findings are instantly available to the Analyst - no copy, no sync→ runnable notebooks in the cookbook, each self-contained and runnable in under 5 minutes
More Recipes
The flagship audit-trail recipe is above. Here are three more common patterns.
from semantica.ingest import FileIngestor
from semantica.split import TextSplitter
from semantica.semantic_extract import NamedEntityRecognizer, RelationExtractor
from semantica.kg import GraphBuilder
from semantica.vector_store import VectorStore, HybridSearch
from semantica.context import AgentContext
# 1. Ingest
docs = FileIngestor().ingest_directory("./docs/", recursive=True)
# 2. Entity-aware chunking - never splits an entity across a chunk boundary
splitter = TextSplitter(method="entity_aware", chunk_size=1000)
chunks = [splitter.split(doc["text"]) for doc in docs]
# 3. Extract entities and relations
ner = NamedEntityRecognizer(confidence_threshold=0.7)
rel_ext = RelationExtractor(confidence_threshold=0.6)
entities = [ner.extract_entities(chunk) for chunk_group in chunks for chunk in chunk_group]
# 4. Build KG
kg = GraphBuilder(merge_entities=True, enable_temporal=True).build(docs)
# 5. Hybrid retrieval
vs = VectorStore(backend="inmemory")
ctx = AgentContext(vector_store=vs, knowledge_graph=kg)
ctx.store("Alice approved the Acme renewal in Q1 2024", conversation_id="c1")
results = HybridSearch(vector_store=vs).search("who approved the renewal?")from semantica.reasoning import ReteEngine, Rule, Fact, RuleType
rete = ReteEngine()
rete.build_network([
Rule(
rule_id="sanctions_check",
name="Flag sanctioned-country transactions",
conditions=[
{"field": "amount", "operator": ">", "value": 10_000},
{"field": "country", "operator": "in", "value": ["IR", "KP", "SY", "CU"]},
],
conclusion="flag_for_compliance_review",
rule_type=RuleType.IMPLICATION,
),
])
# Run the rule across a batch of incoming transactions, not just one
for tx in [
Fact("tx_101", "transaction", [{"amount": 25_000, "country": "IR"}]),
Fact("tx_102", "transaction", [{"amount": 4_500, "country": "DE"}]),
Fact("tx_103", "transaction", [{"amount": 60_000, "country": "KP"}]),
]:
rete.add_fact(tx)
flagged = rete.match_patterns()Same condition-matcher caveat as above applies — validate against your rule set before production use.
from semantica.ingest import FileIngestor
from semantica.semantic_extract import NamedEntityRecognizer, RelationExtractor
from semantica.kg import GraphBuilder
from semantica.ontology import OntologyGenerator, OntologyValidator
from semantica.export import RDFExporter
sources = FileIngestor().ingest_directory("./contracts/")
ner = NamedEntityRecognizer(confidence_threshold=0.7)
entities = ner.process_batch([s["text"] for s in sources])
kg = GraphBuilder(merge_entities=True).build(sources)
gen = OntologyGenerator(base_uri="https://myco.dev/ontology/")
ont = gen.generate_ontology({"entities": entities[0], "relationships": []})
report = OntologyValidator().validate(ont)
if report.valid:
RDFExporter().export({"entities": entities[0]}, "ontology.ttl", format="turtle")Features at a Glance
| Capability | Highlights |
|---|---|
| Context Graphs | Queryable graph of entities, decisions, relationships; causal links; cross-graph navigation |
| Decision Intelligence | record_decision · trace_decision_chain · find_similar_decisions · analyze_decision_impact · check_decision_rules |
| Temporal Intelligence | Point-in-time snapshots · Allen interval algebra (13 relations) · TemporalNormalizer · bi-temporal provenance |
| Distance Intelligence | N×N semantic distance matrices · ego-mode visualization · distance bands · embedding cache |
| Semantic Extraction | NER · relation extraction · event detection · triplet generation · coreference |
| Reasoning Engines | Forward chaining · Rete · deductive · abductive · SPARQL · Datalog with explainable output |
| GraphRAG Chunking | Entity-aware · relation-aware · graph-based · ontology-aware · community-detection chunking |
| Conflict Detection | Value / type / relationship / temporal / logical conflicts · multiple resolution strategies |
| Provenance | W3C PROV-O · every fact traced to source · audit log export JSON/CSV/RDF |
| Ontology Hub | SHACL Studio · visual editor · cross-ontology alignments · health dashboard |
| Vector Store | FAISS · Pinecone · Weaviate · Qdrant · Milvus · PgVector · hybrid + filtered search |
| Graph Databases (LPG) | Neo4j · FalkorDB · Apache AGE · AWS Neptune |
| Triple Stores (RDF) | Oxigraph (embedded) · Blazegraph · Apache Jena · Eclipse RDF4J · unified TripletStore interface · SPARQL query & bulk load |
| Enterprise Data Platforms | Databricks (DatabricksIngestor: Unity Catalog + Delta Lake, PAT/OAuth M2M, table/query ingestion, catalog/schema/table/lineage introspection) · Snowflake (SnowflakeIngestor: warehouse/database/schema, password/key-pair/OAuth auth) |
| LLM Providers | All already supported today: OpenAI (GPT-4o, o1, o3) · Anthropic (Claude) · Google Gemini · Mistral · Meta Llama · Groq · Cohere · Azure OpenAI · AWS Bedrock · Ollama · DeepSeek · Perplexity · Together AI · Fireworks AI · Replicate · HuggingFace · via semantica.llms and LiteLLM |
Performance
Benchmarks from v0.5.0 on a 118,000-node production graph:
| Operation | Before | After | Improvement |
|---|---|---|---|
| Node search (118k nodes) | 24 ms | 0.004 ms | 6,000× faster |
| Embedding cache hit | cold load | revision-based cache | 10× throughput |
| Semantic deduplication | baseline | optimized candidate gen | 6.98× faster |
| Candidate generation | baseline | blocking strategy | 63.6% faster |
Measured on a 118,000-node production graph (AMD EPYC, 64 GB RAM); the deduplication/candidate-generation figures are historical measurements recorded in CHANGELOG.md rather than an automated tests/ assertion. Results vary by hardware, dataset topology, and backend selection — run pytest tests/vector_store/test_performance_benchmarks.py -s to measure your own data.
CLI
Every capability is available from the terminal. The CLI ships with the package, no separate install required.
pip install semantica
semantica # startup dashboard
semantica doctor # health check
semantica --help # full grouped command referenceStart with semantica, verify with doctor, build a graph, and explore the command groups from one terminal.
Command groups: ingest · parse · extract · kg · reason · decision · temporal · provenance · ontology · embed · deduplicate · validate · export · visualize · pipeline · server · explorer · mcp · doctor · shell · init · watch
Integrations
Native plugin bundles for Claude Code, Cursor, Codex, Windsurf, Cline, Continue, VS Code, and OpenClaw; a full-featured MCP server for any MCP-compatible client; a comprehensive REST API; and first-class Agno and CrewAI support for agentic frameworks. Every major LLM provider is already supported via semantica.llms and LiteLLM: OpenAI, Anthropic, Gemini, Mistral, Llama, Groq, Cohere, Azure, Bedrock, Ollama, DeepSeek, HuggingFace, and more.
MCP setup takes 30 seconds — see MCP Server below.















Agentic Frameworks













MCP Server
Connect any MCP-compatible client (Claude Desktop, Windsurf, Cline, VS Code) in 30 seconds:
python -m semantica.mcp_server
# or via the installed entry point
semantica-mcp{
"mcpServers": {
"semantica": { "command": "python", "args": ["-m", "semantica.mcp_server"] }
}
}Tools exposed over MCP:
| Tool | What it does |
|---|---|
extract_entities | NER on any text |
extract_relations | Relation extraction |
record_decision | Persist a decision node |
query_decisions | Search decision history |
find_precedents | Semantic precedent lookup |
get_causal_chain | Full causal ancestry |
add_entity | Add a KG node |
add_relationship | Add a KG edge |
run_reasoning | Execute rule set |
get_graph_analytics | Centrality, communities |
export_graph | Export to RDF/JSON/Parquet |
get_graph_summary | Graph statistics |
REST API
# Start the backend
python -m semantica.server # port 8000
# Extract entities & relations via REST
curl -X POST http://localhost:8000/api/enrich/extract \
-H "Content-Type: application/json" \
-d '{"text": "Apple CEO Tim Cook announced record earnings."}'
# List recorded decisions
curl "http://localhost:8000/api/decisions?category=vendor_selection"
# Query the knowledge graph
curl "http://localhost:8000/api/graph/node/acme_corp/neighbors?depth=2"REST endpoints span: enrich (extract) · graph · decisions · reasoning · provenance · ontology · embeddings · search · export · pipeline · temporal · deduplication
Plugin Bundles
Domain skills: extract · ingest · query · ontology · validate · deduplicate · embed · reason · decision · causal · temporal · provenance · policy · explain · export · change · visualize
Specialized agents: kg-assistant · decision-advisor · explainability
Bundles for Claude Code, Cursor, Codex, Windsurf, Cline, Continue, VS Code, and OpenClaw in plugins/.
Knowledge Explorer
A browser-based graph workbench. Pan and zoom live graphs, scrub the timeline, review every decision's causal chain, resolve duplicates, and author your ontology visually. Built on React 19 + Sigma.js.
| Workspace | What you can do |
|---|---|
| Knowledge Graph | Live Sigma.js canvas with ForceAtlas2 layout, Ego Mode, semantic distance heatmap |
| Timeline | Scrub through temporal events and watch the graph evolve |
| Decisions | Browse the causal chain behind every recorded decision |
| Registry | Live audit log of every graph mutation |
| Entity Resolution | Review and merge duplicates |
| Ontology Hub | SHACL Studio, visual editor, cross-ontology alignments, SKOS browser |
| Lineage | W3C PROV-O provenance visualization for any entity |
Quickest way to start (no Node.js required):
pip install "semantica[explorer]"
semantica-explorer --graph my_graph.json
# Dashboard opens at http://127.0.0.1:8000For contributor / dev-server setup: explorer/README.md: Local Setup Guide
What's New in v0.6.5
Security release — upgrading is strongly recommended. Fixes for 5 externally-reported vulnerabilities in the Explorer API and graph/triplet store backends, plus a CodeQL-flagged ReDoS:
- Missing authentication on all Explorer API routes (GHSA-j4mq-hprp-987v, Critical): every route now requires
SEMANTICA_API_KEY, fails closed (503) rather than open when unconfigured - SSRF via redirect bypass in ontology URL fetching (GHSA-8c7v-62gr-hj6g, High): redirect targets are now re-validated at every hop and the connection is pinned to the validated address, closing a DNS check-then-use race
- Cypher injection via unvalidated node labels and property keys (GHSA-482h-hw99-h62p, Critical): Neptune, Neo4j, and FalkorDB now sanitize every label/relationship-type/property-key interpolation site
- SPARQL injection via unvalidated triplet IRIs (GHSA-8vgg-8mr4-r236, Critical): Blazegraph, RDF4J, and Jena now validate subject/predicate/object IRIs before interpolation
- Missing Origin validation on the WebSocket handshake (GHSA-4643-wpgq-w329, Moderate, anonymous-mode only):
/ws/graph-updatesnow checksOriginagainst the same allowlistCORSMiddlewareenforces for HTTP - Polynomial ReDoS in SPARQL query validation (CodeQL
py/polynomial-redos): fixed a backtracking regex in the Explorer's SPARQL route
Also includes: embedded Oxigraph backend for TripletStore, PROV-O trust/spec completeness for ProvenanceManager, and the Altair Anzo triplet store backend.
→ Full release notes · Changelog
Built for High-Stakes Domains
Semantica is designed for environments where AI outputs must be explainable, auditable, and defensible, and where the data itself can't leave your infrastructure. Self-hostable with zero vendor lock-in, it's built as much for organizations handling confidential or classified data as for regulated industries chasing an audit trail:
- Finance: Loan underwriting audit trails, fraud detection, AML compliance, regulatory risk knowledge graphs
- Healthcare: Clinical decision support, drug interaction graphs, and patient safety audit trails
- Legal: Evidence-backed research, contract analysis, case law reasoning, and privilege tracking
- Government & Defense: Policy decision records, classified information governance, and regulatory reporting, fully self-hosted with no data leaving your perimeter
- Law Enforcement: Case linkage, evidence provenance chains, and investigative knowledge graphs that hold up under legal scrutiny
- Cybersecurity: Threat attribution, incident response timelines, and IOC provenance tracking
- Autonomous Systems: Decision logs, safety validation, and explainable AI for certification
⚠️ This is system-level explainability, not foundation-model explainability. Semantica does not expose, reconstruct, or explain what happens inside the LLM/foundation model — its internal reasoning or chain-of-thought stays opaque, as it does for any external system. What Semantica explains is outside the model: the context and data fed in, the decision produced, its provenance, the relevant relationships, the policies applied, and the full execution trail. In short, Semantica explains and audits what the AI system did, not the LLM's private internal reasoning.
Installation
pip install semantica # core
pip install semantica[all] # everythingpip install semantica[agno] # Agno multi-agent integration
pip install semantica[crewai] # CrewAI integration
pip install semantica[llm-litellm] # OpenAI, Anthropic, Gemini, Mistral, Llama, Groq, Cohere, Bedrock, Ollama, DeepSeek, and more
pip install semantica[graph-neo4j] # Neo4j graph store (LPG)
pip install semantica[graph-falkordb] # FalkorDB graph store (LPG)
pip install semantica[graph-apache-age] # Apache AGE graph store (LPG)
pip install semantica[graph-amazon-neptune] # AWS Neptune graph store (LPG)
pip install semantica[tripletstore-oxigraph] # Embedded in-memory/on-disk RDF store
# RDF triple stores (Blazegraph, Apache Jena, Eclipse RDF4J) need no extra:
# semantica.triplet_store talks SPARQL over HTTP using the core `requests` dependency
pip install semantica[vectorstore-qdrant] # Qdrant vector store
pip install semantica[vectorstore-pinecone] # Pinecone vector store
pip install semantica[db-snowflake] # Snowflake
pip install semantica[db-databricks] # Databricks (SDK + SQL connector)
pip install semantica[ingest-parquet] # Parquet / PyArrow
pip install semantica[ingest-arrow] # Apache Arrow, Feather, IPC
pip install semantica[viz] # HTML interactive visualization
pip install semantica[watch] # Directory file watcher
pip install semantica[explorer] # Knowledge Explorer dashboardFor production deployments, use Docker or Kubernetes rather than a local pip install. Set SEMANTICA_SECRET_KEY, configure a persistent LPG graph store (Neo4j / FalkorDB / Apache AGE / AWS Neptune) and/or RDF triple store (Blazegraph / Apache Jena / Eclipse RDF4J), and point the vector store at a hosted backend (Qdrant / Pinecone). See ARCHITECTURE.md for the full deployment topology.
# From source
git clone https://github.com/semantica-agi/semantica.git
cd semantica && pip install -e ".[dev]" && pytest tests/Enterprise
On-premises deployment · Private cloud · Custom domain implementations · SLA-backed support · Professional services for regulated industries (finance, healthcare, legal, government).
getsemantica.ai for enterprise solutions and pricing.
Community & Support
| Discord | discord.gg/sV34vps5hH: real-time help, showcases, and announcements |
| GitHub Discussions | Q&A and feature requests |
| GitHub Issues | Bug reports |
| Documentation | docs.getsemantica.ai |
| Cookbook | Runnable Jupyter notebooks |
| Changelog | CHANGELOG.md · Release Notes |
Star History
Contributors
Contributing
All contributions are welcome: bug fixes, features, tests, and documentation.
- 1Fork the repo and create a branch
- 2
pip install -e ".[dev]" - 3Write tests alongside your changes (
pytest tests/) - 4Open a PR and tag
@KaifAhmad1for review
See CONTRIBUTING.md for full guidelines.
MIT License · Built by Semantica
GitHub · Discord · Twitter/X · Website · Docs · PyPI
If this project helps you build better AI, a star means a lot.
English · Deutsch · Français · Español · Italiano · Português · العربية · اردو · हिन्दी · 中文 · 日本語 · 한국어
