#textsummarisation #bert #naturallanguageprocessing
This video talks about very recent work in the domain of text summarization. Researchers introduce a method for generating extended extractive summaries of Long documents such as research papers, legal document using Multi-task loss. The work surpasses BERTSUM on ROUGE scale.
P.S. Sorry for bad audio at certain segments in the video.
⏩ Abstract: Prior work in document summarization has mainly focused on generating short summaries of a document. While this type of summary helps get a high-level view of a given document, it is desirable in some cases to know more detailed information about its salient points that can't fit in a short summary. This is typically the case for longer documents such as a research paper, legal document, or a book. In this paper, we present a new method for generating extended summaries of long papers. Our method exploits hierarchical structure of the documents and incorporates it into an extractive summarization model through a multi-task learning approach. We then present our results on three long summarization datasets, arXiv-Long, PubMed-Long, and Longsumm. Our method outperforms or matches the performance of strong baselines. Furthermore, we perform a comprehensive analysis over the generated results, shedding insights on future research for long-form summary generation task. Our analysis shows that our multi-tasking approach can adjust extraction probability distribution to the favor of summary-worthy sentences across diverse sections.
Please feel free to share out the content and subscribe to my channel :)
⏩ Subscribe - / @techvizthedatascienceguy
⏩ OUTLINE:
0:00 - Abstract
3:42 - Dataset (Longsumm, arXiv-long, PubMed-long)
4:54 - BERTSUM Paper Explained (Extractive Summarization Method)
10:00 - Section Aware Summarizer
12:26 - Analysis and Results
⏩ Paper Title: On Generating Extended Summaries of Long Documents
⏩ Paper: https://arxiv.org/pdf/2012.14136.pdf
⏩ Paper Code: https://github.com/Georgetown-IR-Lab/...
⏩ Author: Sajad Sotudeh, Arman Cohan, Nazli Goharian
⏩ Organisation: IR Lab, Georgetown University, Allen Institute for Artificial,
⏩ IMPORTANT LINKS
BERTSUM Paper - https://arxiv.org/abs/1903.10318
BERT for Summarizing Lectures - • Leveraging BERT for Extractive Text Summar...
COVID-19 Article Summarization - • Text Summarization of COVID-19 Medical Art...
Extractive and Abstractive Summarization using Transformer Language Models - • Extractive & Abstractive Summarization wit...
PEGASUS: Abstractive Summarization - • PEGASUS: Pre-training with Gap-Sentences f...
*********************************************
If you want to support me financially which totally optional and voluntary :) ❤️
You can consider buying me chai ( because i don't drink coffee :) ) at https://www.buymeacoffee.com/TechvizC...
*********************************************
⏩ Youtube - / @techvizthedatascienceguy
⏩ Blog - https://prakhartechviz.blogspot.com
⏩ LinkedIn - / prakhar21
⏩ Medium - / prakhar.mishra
⏩ GitHub - https://github.com/prakhar21
*********************************************
Please feel free to share out the content and subscribe to my channel :)
⏩ Subscribe - / @techvizthedatascienceguy
Tools I use for making videos :)
⏩ iPad - https://tinyurl.com/y39p6pwc
⏩ Apple Pencil - https://tinyurl.com/y5rk8txn
⏩ GoodNotes - https://tinyurl.com/y627cfsa
#techviz #datascienceguy #machinelearning #nlproc #researchpaper #explained #transformers