Summarize web articles: extract clean content, title, image, language, and generate text summaries
-
Updated
Jul 19, 2026 - Python
Summarize web articles: extract clean content, title, image, language, and generate text summaries
This repository is part of an NLP course for humanities and cultural studies. This course uses historical newspapers as a source and applies NLP methods to them. NLP tasks: Tokenization, Lemmatization, TF-IDF, Part-of-speech tagging, semantic search with transformers, article extraction and OCR post-correction with LLMs, NER and text classification
GNewsScraper is a TypeScript package that scrapes article data from Google News based on a keyword or phrase. It returns the results as an array of JSON objects, making it convenient to access and use the scraped information
A simple HTML-to-Markdown converter with article extraction, selector filtering, and batch conversion.
强大易用的文章转 Markdown 工具,一键采集微信公众号、掘金、CSDN 等文章,自动下载图片、清理代码块、支持批量转换和 Agent Skill
OpenClaw Skill:读取微信公众号文章、识别公众号并拉取文章列表
📋 WebMD is a Chrome extension that transforms web pages into Markdown documents with surgical precision.
A pipe-based news article scraping and metadata extraction library for Python
A configurable pipeline for extracting and filtering articles from large corpora, tailored for the Delpher Kranten corpus, with support for features like keyword filtering and tf-idf-based relevance scoring.
Mantis grabs exactly what you'd see on the page and turns it into structured JSON or clean Markdown. I built it for read-later and bookmarking tools, and for AI agents that need input without all the token bloat. It's Readability-style, has zero dependencies, and lives in a single file.
Zendesk articles extraction toolkit
Chrome/Edge extension that estimates article word count, reading time, and lets you double-click any word to update the toolbar badge with your reading progress.
Chrome extension that yoinks webpages into clean markdown. Supports article extraction, full-page capture, YouTube transcripts, and visual element picking.
HTML main-content extraction for Rust — ports of Mozilla Readability, Trafilatura, and htmldate.
Production web scraper with Playwright, bot-detection plugins, fingerprint rotation, and CAPTCHA solving. CLI + FastAPI.
KSL news article scraper
India Times news extraction tool
US Magazine news extractor
Daily Beast news scraper
Add a description, image, and links to the article-extraction topic page so that developers can more easily learn about it.
To associate your repository with the article-extraction topic, visit your repo's landing page and select "manage topics."