CorpusMind Lens

CorpusMind Lens

Vision-LM-Powered Multimodal Discourse Analysis

تحليل الخطاب متعدد الأنماط المدعوم بنماذج الرؤية اللغوية

CorpusMind Lens

كوربَس مايند لِنز

Dr. Waleed Mandour (Sultan Qaboos University) · Prof. Wesam Ibrahim (Princess Nourah Bint Abdulrahman University) د. وليد مندور (جامعة السلطان قابوس) · أ.د. وسام إبراهيم (جامعة الأميرة نورة بنت عبدالرحمن)

A standalone desktop application for vision-language-model-powered image understanding and multimodal discourse analysis

Vision-LMنموذج الرؤية 8 Frameworks٨ أطر Alignmentالمحاذاة Consent Gateبوابة الموافقة Arabic + Englishالعربية + الإنجليزية Local-Firstمحلي أولاً Free & Open Sourceمجاني ومفتوح المصدر

CorpusMind Lens is a companion application to CorpusMind, designed specifically for researchers working at the intersection of visual and linguistic discourse. It connects to the same CorpusMind engine but provides a focused interface for image-set management, vision-language-model image description, and multimodal discourse analysis across eight theoretical frameworks. The application operates entirely on local hardware using Ollama or LM Studio — no image data, model queries, or analytical output is transmitted to any external server. All interpretive claims are phrased as hypotheses grounded in specific visual or textual evidence, following the methodological commitment to transparency and reproducibility that underpins the CorpusMind platform.

Analytical Capabilities

Image-Set Management

إدارة مجموعات الصور

Researchers organise images into sets within a corpus, upload via drag-and-drop with optional captions, and receive automatic cached analysis including Tesseract OCR text extraction, dominant colour palette computation, and composition geometry (rule-of-thirds, salience centre, visual balance).

يُنظّم الباحثون الصور في مجموعات ضمن المدونة، ويُرفعونها بالسحب والإفلات مع تعليقات اختيارية، ويتلقّون تحليلًا تلقائيًا مخزّنًا يشمل استخراج النصوص عبر Tesseract، وحساب لوحة الألوان السائدة، وهندسة التكوين (قاعدة الأثلاث، مركز البروز، التوازن البصري).

Vision-Language-Model Description

الوصف عبر نموذج الرؤية اللغوية

Image bytes are transmitted directly to a local vision-language model (e.g., moondream, llama3.2-vision) which produces a grounded description. The system records full provenance metadata — model name, provider, prompt hash, timestamp — for reproducibility. When the model's text reading diverges from the cached Tesseract OCR, both are presented side by side rather than silently overwriting either.

تُرسل وحدات بايت الصورة مباشرةً إلى نموذج رؤية لغوية محلي (مثل moondream أو llama3.2-vision) الذي يُنتج وصفًا مُستندًا إلى الأدلة. يُسجّل النظام بيانات المصدر الكاملة — اسم النموذج، المُزوّد، تجزئة الاستفسار، الطابع الزمني — لضمان إعادة الإنتاج. عندما تختلف قراءة النموذج عن نص Tesseract المخزّن، يُعرض كلاهما جنبًا إلى جنب بدلاً من الكتابة فوق أحدهما بصمت.

Eight Discourse-Theoretical Frameworks

ثمانية أطر نظرية لتحليل الخطاب

Each framework operates in two modes: a fast heuristic path (colour, composition, OCR, caption) and a vision-LM path (?mode=llm) that sends the image and the framework's theoretical lens to the model. Frameworks: Social Semiotic (Kress & van Leeuwen), CDA (Fairclough, van Dijk, Wodak, Machin & Mayr), Persuasion (Aristotle + Toulmin), Framing (Entman), Narrative (Labov), Visual Metaphor (MIPVU-inspired), Emotion, and Cultural analysis.

يعمل كل إطار في وضعين: مسار استدلالي سريع (اللون، التكوين، استخراج النص، التعليق) ومسار نموذج الرؤية (?mode=llm) الذي يُرسل الصورة وعدسة الإطار النظري إلى النموذج. الأطر: التضمين الاجتماعي (كريس وفان ليوين)، التحليل النقدي للخطاب (فيركلوف، فان دايك، فوداك، ماشين وماير)، الإقناع (أرسطو + تولمين)، التأطير (إنتمان)، السرد (لابوف)، الاستعارة البصرية (مُستوحى من MIPVU)، العاطفة، والتحليل الثقافي.

Image-Text Alignment Inspector

مُفتّش محاذاة الصورة والنص

Aligns image regions with text spans using three modes: heuristic (colour-term and positional matching), vision-LM (the model identifies which text fragments refer to which image regions), or both displayed side by side. The user can compare which mode produced which alignment — maintaining analytical transparency.

يُحاذي مناطق الصورة مع مقاطع النص باستخدام ثلاثة أوضاع: استدلالي (مطابقة المصطلحات اللونية والموضعية)، ونموذج الرؤية (يُحدد النموذج أي المقاطع النصية تشير إلى أي مناطق في الصورة)، أو كلاهما معروضًا جنبًا إلى جنب. يمكن للمستخدم مُقارنة أي وضع أنتج أي محاذاة — مما يُحافظ على الشفافية التحليلية.

Cross-Image Batch Analysis

التحليل الدفعي عبر الصور

Aggregates cached vision-LM analysis across all images in a set: surfaces recurring framework themes (grouped by theoretical lens, counted by category), produces an OCR-derived word-frequency list, and summarises all vision-LM descriptions. Enables pattern identification across a corpus of images without manual inspection of each.

يُجمّع التحليل المخزّن بنموذج الرؤية عبر جميع الصور في المجموعة: يُبرز السمات المتكررة للأطر (مُجمّعة حسب العدسة النظرية، مَعدودة حسب الفئة)، ويُنتج قائمة تردد كلمات مستمدة من استخراج النصوص، ويُلخّص جميع أوصاف نموذج الرؤية. يُتيح تحديد الأنماط عبر مدونة من الصور دون فحص يدوي لكل صورة.

Ethical Consent Gate

بوابة الموافقة الأخلاقية

All person-descriptive content generated by vision-language models — including age, gender, facial expression, physical appearance, ethnicity, religious attire, and socioeconomic indicators — is automatically redacted from responses unless the researcher explicitly enables facial analysis in Settings. The redaction operates at the sentence level and applies across all vision-LM routes, including cached results.

يُحجب تلقائيًا جميع المحتوى الوصفي للأشخاص المُولّد بواسطة نماذج الرؤية اللغوية — بما في ذلك العمر، والجنس، وتعابير الوجه، والمظهر الجسدي، والعرق، والملابس الدينية، والمؤشرات الاجتماعية والاقتصادية — من الردود ما لم يُفعّل الباحث تحليل الوجه صراحةً في الإعدادات. يعمل الحجب على مستوى الجملة ويُطبّق عبر جميع مسارات نموذج الرؤية، بما في ذلك النتائج المخزّنة.

How CorpusMind Lens Compares

The following table compares CorpusMind Lens with established tools in qualitative data analysis, multimodal annotation, and cloud-based computer vision. Each tool has its strengths; CorpusMind Lens is the only one that combines vision-language models with explicit discourse-theoretical frameworks, local-first deployment, first-class Arabic support, and a reproducible analytical pipeline — all in a single free, open-source application.

يُقارن الجدول التالي كوربَس مايند لِنز مع الأدوات المعتمدة في تحليل البيانات النوعية، والوسم متعدد الأنماط، والرؤية الحاسوبية السحابية. لكل أداة نقاط قوتها؛ غير أنّ كوربَس مايند لِنز هو الأداة الوحيدة التي تجمع بين نماذج الرؤية اللغوية والأطر النظرية الصريحة لتحليل الخطاب، والتشغيل المحلي، والدعم المتقدم للغة العربية، ومسار تحليلي قابل لإعادة الإنتاج — في تطبيق واحد مجاني ومفتوح المصدر.

Capabilityالقدرة CorpusMind LensNVivoATLAS.tiMAXQDAELAN Google / Azure CVGoogle / Azure CV Multimodal Analysis (O'Halloran)التحليل متعدد الأنماط (أوهالوران)
Multimodal discourse analysis (image + text)تحليل الخطاب متعدد الأنماط (صورة + نص)~~~~
Visual grammar (Kress & van Leeuwen)القواعد البصرية (كريس وفان ليوين)~
Critical discourse analysis (images)التحليل النقدي للخطاب (الصور)
Image–text alignmentمحاذاة الصورة والنص~
Vision-language model integrationتكامل نموذج الرؤية اللغوية~~~
Reproducible / auditable pipelineمسار تحليلي قابل لإعادة الإنتاج والتدقيق~~~~
Local / offline capableيعمل محليًا / دون اتصال~~~
Arabic support (NLP + vision)دعم اللغة العربية (معالجة + رؤية)~~~~~
Free / open-sourceمجاني / مفتوح المصدر~
Privacy: data stays on deviceالخصوصية: تبقى البيانات على الجهاز~~
Ethical consent gate for person descriptionsبوابة موافقة أخلاقية للأوصاف الشخصية

Full support  ·  ~ Partial / limited  ·  Not available. Cloud vision APIs (Google Cloud Vision, Azure Computer Vision) provide VLMs and Arabic OCR but lack discourse frameworks, visual grammar, image-text alignment, reproducible pipelines, and local deployment.

دعم كامل  ·  ~ جزئي / محدود  ·  غير متاح. توفر واجهات الرؤية السحابية (Google Cloud Vision وAzure Computer Vision) نماذج رؤية لغوية واستخراج نصوص عربي، لكنها تفتقر إلى أطر تحليل الخطاب، والقواعد البصرية، ومحاذاة الصورة والنص، والمسارات القابلة لإعادة الإنتاج، والتشغيل المحلي.

Vision-Language Models

CorpusMind Lens supports any vision-capable model available through Ollama or LM Studio. The following models are listed in the in-app download catalogue (Settings → Model Providers) and can be installed directly from within the application:

يدعم كوربَس مايند لِنز أي نموذج رؤية متاح عبر Ollama أو LM Studio. النماذج التالية مُدرجة في كتالوج التنزيل داخل التطبيق (الإعدادات ← مُزوّدو النماذج) ويمكن تثبيتها مباشرةً من داخل التطبيق:

moondream

1.7 GB · 1.8B parameters · 4 GB RAM
Small vision-language model. Runs on any machine. Suitable for basic image description and alignment tasks.

‏1.7 جيجابايت · ‏1.8 مليار مَعامل · ‏4 جيجابايت رام
نموذج رؤية لغوية صغير. يعمل على أي جهاز. مناسب لمهام وصف الصور والمحاذاة الأساسية.

Ollama Library →

llama3.2-vision:11b

7.9 GB · 11B parameters · 8 GB RAM
Higher-quality image understanding. Superior for complex visual discourse analysis. Requires 8 GB+ RAM.

‏7.9 جيجابايت · ‏11 مليار مَعامل · ‏8 جيجابايت رام
فهم عالي الجودة للصور. متفوق في تحليل الخطاب البصري المعقد. يتطلب ‏8 جيجابايت رام أو أكثر.

Ollama Library →

gemma3:4b

3.3 GB · 4B parameters · 4 GB RAM
Multilingual (including Arabic). Supports vision in addition to text. Good balance of size and capability.

‏3.3 جيجابايت · ‏4 مليارات مَعامل · ‏4 جيجابايت رام
متعدد اللغات (بما في ذلك العربية). يدعم الرؤية بالإضافة إلى النص. توازن جيد بين الحجم والقدرة.

Ollama Library →

Download CorpusMind Lens

Version 1.0.0 — free, open-source (AGPL-3.0-only). No subscription, no telemetry, no cloud dependency.

Version 1.0.0 · AGPL-3.0-only · Third-party licenses

Live Statistics

Lens Downloadsتحميلات لِنز
GitHub Starsنجوم GitHub
Lens Assetsملفات لِنز
Welcome! أهلاً بك!
You are visitor أنت الزائر رقم
Detecting… جارٍ التحديد…
Your location موقعك

Downloads by Platformالتحميلات حسب المنصة

macOS
0
Windows
0
Linux
0
User Guideالدليل
0

Data fetched from GitHub Releases API · updates in real timeالبيانات من واجهة GitHub Releases · تُحدّث في الوقت الفعلي