تحويل PDF إلى نص — المهندس سعيد باعطية
استخرج نص أي PDF داخل متصفحك
عن الأداة
استخرج نص أي PDF داخل متصفحك
استخرج النص من ملفات PDF مباشرة داخل متصفحك دون رفعها إلى أي خادم. يكشف الأعمدة والجداول، ويدعم النصوص العربية مع فكّ ترميز خطوط CID، ويحوّل العناوين والقوائم والجداول إلى Markdown منسّق. يكشف تلقائياً الملفات الممسوحة ضوئياً ويخبرك بالصفحات التي لا تحتوي نصاً قابلاً للاستخراج، ويفتح الملفات المحمية بكلمة مرور.
- السعر
- مجانية بالكامل
- مكان المعالجة
- داخل متصفحك بالكامل — لا يُرفع أي ملف أو بيانات إلى أي خادم
- متطلبات التشغيل
- تحتاج هذه الأداة إلى JavaScript لتعمل، لأن كل الحسابات تجري على جهازك.
- المطوّر
- المهندس سعيد باعطية
طريقة الاستخدام
- 1. أفلت ملف PDF — اسحب الملف إلى المنطقة المخصصة أو اخترْه من جهازك. الملف يُقرأ داخل متصفحك ولا يُرفع إلى أي خادم.
- 2. اضبط خيارات الاستخراج — اختر نمط المخرجات (مطابق للمصدر أو مختصر)، وحدّد نطاق صفحات إن أردت، وأدخل كلمة المرور إن كان الملف محمياً.
- 3. راجع تقرير المستند — اطّلع على نوع الملف ونسبة الثقة وعدد الصفحات، وعلى أي تنبيه بصفحات ممسوحة ضوئياً لا تحتوي نصاً.
- 4. انسخ النتيجة أو نزّلها — بدّل بين مخرجات Markdown والنص العادي، ثم انسخ المحتوى أو نزّله بصيغة .md أو .txt.
قدرات الأداة
- نص مستخرج داخل متصفحك بالكامل — الملف لا يغادر جهازك ولا يُرفع إلى أي خادم
- كشف الأعمدة والجداول، فتخرج كتلاً منفصلة بدل تسلسل عشوائي للسطور
- معالجة النص العربي وفكّ ترميز خطوط CID عبر جداول ToUnicode
- مخرجات Markdown منسّقة تحفظ العناوين والقوائم والجداول وكتل الكود
- كشف تلقائي لنوع المستند مع تحديد أرقام الصفحات التي لا تحتوي نصاً قابلاً للاستخراج
- تنبيه صريح عند رصد مشاكل ترميز في خطوط المستند
- تنزيل بصيغتَي .md و.txt أو نسخ مباشر إلى الحافظة
حدود الأداة
الأداة تقرأ النص المخزّن فعلياً داخل ملف PDF، وهذا يغطي كل مستند أُنشئ من Word أو Google Docs أو نظام فوترة أو أي برنامج. ما لا تفعله هو التعرّف الضوئي على الحروف: إن كان الملف صوراً ممسوحة بماسح ضوئي أو صور جوال، فلا نص بداخله ليُقرأ. في هذه الحالة تخبرك الأداة صراحة بنوع الملف وبأرقام الصفحات المتأثرة بدل أن تعطيك ناتجاً فارغاً بلا تفسير.
أسئلة شائعة
- هل يُرفع ملف PDF إلى خادم؟
- لا. المعالجة كلها تجري داخل متصفحك عبر وحدة WebAssembly تُحمَّل مرة واحدة، وبايتات الملف لا تُرسل إلى أي مكان. لهذا لا يوجد حد لحجم الملف تفرضه أنا، ولهذا تستطيع استخدام الأداة على مستندات حساسة كالعقود والسجلات. للتحقق: افتح تبويب الشبكة في أدوات المطور، أفلت ملفك، وستجد أن الطلب الوحيد هو تحميل ملف المعالجة نفسه — ولا طلب يحمل ملفك.
- هل يدعم ملفات PDF العربية؟
- نعم، والمحرّك يفكّ ترميز خطوط CID عبر جداول ToUnicode، وهي الآلية التي تنكسر عندها أغلب الأدوات فتُخرج العربية كرموز مشوّهة. لكن جودة الناتج تعتمد على البرنامج الذي أنشأ الملف: المستندات الصادرة من Word أو InDesign أو أنظمة التقارير تخرج نظيفة، بينما بعض ملفات PDF المطبوعة من المتصفح تخزّن الحروف بترتيبها البصري لا المنطقي، فتظهر كلمات فيها لام-ألف أو همزة بترتيب حروف مختلّ. هذا قيد في تلك الملفات نفسها لا في الأداة — وتُنتج المكتبات المرجعية مثل pdf.js النتيجة ذاتها عليها. الأداة تنبّهك صراحةً عند رصد مشاكل ترميز في خطوط المستند.
- رفعت ملفاً وظهرت رسالة أنه ممسوح ضوئياً ولم يخرج نص — لماذا؟
- لأن الملف لا يحتوي نصاً أصلاً، بل صوراً لصفحات. هذا يحدث عندما يكون المستند ممسوحاً بماسح ضوئي أو مصوّراً بالجوال: ما تراه حروفاً هو في الحقيقة بكسلات في صورة. استخراج النص منه يتطلب تعرّفاً ضوئياً على الحروف (OCR) وهو عملية مختلفة غير مدعومة في هذه الأداة. الأداة تكشف الحالة وتخبرك بأرقام الصفحات المتأثرة بدل أن تتركك أمام نتيجة فارغة.
- ما الفرق بين مخرجات Markdown والنص العادي؟
- النص العادي يعطيك الحروف فقط بلا أي بنية. أما Markdown فيحفظ شكل المستند: العناوين تبقى عناوين، والقوائم تبقى قوائم، والجداول تبقى جداول. اختر Markdown إن كنت ستلقّم النص لنموذج ذكاء اصطناعي مثل ChatGPT أو Claude — فالبنية تساعده على فهم المستند بدقة أعلى — أو إن كنت ستنقل المحتوى إلى محرر أو نظام إدارة محتوى. واختر النص العادي إن كنت تريد الكلمات وحدها للبحث أو المعالجة.
- هل يفتح ملفات PDF المحمية بكلمة مرور؟
- نعم، إن كنت تعرف كلمة المرور. يوجد حقل لإدخالها، وتُستخدم داخل متصفحك فقط لفك تشفير الملف أثناء القراءة ولا تُحفظ ولا تُرسل. الأداة لا تكسر الحماية ولا تتجاوز كلمة مرور لا تملكها.
خدمة مرتبطة
ربط الأنظمة معاً بسلاسة
التواصل المباشر
نموذج التواصل في الموقع يحتاج JavaScript. القنوات التالية تعمل بدونه تمامًا.
PDF to Text Converter — Engineer Saeed Baatiah
Extract any PDF’s text in your browser
About this tool
Extract any PDF’s text in your browser
Extract text from PDF files directly in your browser without uploading them to any server. It detects columns and tables, handles Arabic text with CID font decoding, and converts headings, lists and tables into formatted Markdown. It automatically detects scanned files and tells you which pages hold no extractable text, and it opens password-protected files.
- Price
- Completely free
- Where processing happens
- Entirely inside your browser — no file or data is uploaded to any server
- Runtime requirement
- This tool needs JavaScript to run, because all of the computation happens on your own device.
- Built by
- Engineer Saeed Baatiah
How to use it
- 1. Drop a PDF — Drag the file onto the drop zone or pick it from your device. It is read inside your browser and never uploaded.
- 2. Set extraction options — Choose the output profile (faithful or compact), set a page range if you want one, and enter a password if the file is protected.
- 3. Read the document report — Check the document type, confidence, page count, and any notice about scanned pages that hold no text.
- 4. Copy or download the result — Switch between the Markdown and plain text output, then copy the content or download it as .md or .txt.
Tool capabilities
- Text extracted entirely inside your browser — the file never leaves your device
- Columns and tables detected, so you get separate blocks rather than interleaved lines
- Arabic text handling with CID font decoding through ToUnicode tables
- Formatted Markdown output preserving headings, lists, tables and code blocks
- Automatic document-type detection naming the pages that hold no extractable text
- An explicit warning when the document’s fonts show encoding problems
- Download as .md or .txt, or copy straight to the clipboard
Limits and scope
The tool reads text that is genuinely stored inside the PDF, which covers every document produced by Word, Google Docs, a billing system, or any other program. What it does not do is optical character recognition: if the file is scanner or phone images, there is no text inside it to read. In that case the tool tells you the document type and the affected page numbers outright, rather than handing you an empty result with no explanation.
Frequently asked questions
- Is my PDF uploaded to a server?
- No. All processing happens inside your browser through a WebAssembly module loaded once, and the file’s bytes are never sent anywhere. That is why there is no file-size limit imposed by me, and why you can safely use it on sensitive documents like contracts and records. To verify: open the network tab in your developer tools, drop your file, and you will see the only request is for the processing module itself — none carrying your file.
- Does it support Arabic PDF files?
- Yes, and the engine decodes CID fonts through their ToUnicode tables — the exact mechanism where most tools break and spit out Arabic as mangled symbols. Output quality does depend on what produced the file, though: documents from Word, InDesign or reporting systems come out clean, while some browser-printed PDFs store letters in visual rather than logical order, so words containing lam-alef or hamza can come back with their letters out of sequence. That is a limitation of those files rather than of the tool — reference libraries such as pdf.js return the same result on them. The tool tells you outright when it detects encoding problems in a document’s fonts.
- I uploaded a file and got a “scanned” message with no text — why?
- Because the file contains no text at all, only page images. That happens when a document was run through a scanner or photographed with a phone: what looks like letters is really pixels in an image. Pulling text out of that needs optical character recognition (OCR), a different process this tool does not perform. It detects the situation and names the affected pages instead of leaving you with a blank result.
- What is the difference between the Markdown and plain text output?
- Plain text gives you the characters and nothing else. Markdown preserves the document’s shape: headings stay headings, lists stay lists, tables stay tables. Choose Markdown if you are feeding the text to an AI model like ChatGPT or Claude — the structure helps it read the document more accurately — or if you are moving the content into an editor or CMS. Choose plain text when you just want the words for searching or processing.
- Does it open password-protected PDF files?
- Yes, if you know the password. There is a field to enter it, and it is used inside your browser only to decrypt the file while reading, never stored and never transmitted. The tool does not break protection or bypass a password you do not have.
Related paid service
Connect systems seamlessly
Direct contact
The site contact form needs JavaScript. Every channel below works without it.
This page in another language