# Large Language Models — science-database.com > LLM research and development. Transformer architectures, training methods, alignment, reasoning capabilities, multimodal models, and AI safety. - Discipline: Computer Science / AI - URL: https://science-database.com/technology/large-language-models - API: https://science-database.com/api/v1/technology/large-language-models - Last Updated: 2026-09-08T05:36:38.035Z - Articles Indexed: 15 ## Top Publications ### Gradient-based learning applied to document recognition - Authors: Yann LeCun, Léon Bottou, Yoshua Bengio, Patrick Haffner - Journal: Proceedings of the IEEE - Date: 1998-01-01 - DOI: https://doi.org/10.1109/5.726791 - Citations: 59536 - Source: OpenAlex - llms.txt: https://science-database.com/technology/large-language-models/paper/oa-W2112796928/llms.txt - Abstract: Multilayer neural networks trained with the back-propagation algorithm constitute the best example of a successful gradient based learning technique. Given an appropriate network architecture, gradient-based learning algorithms can be used to synthesize a complex decision surface that can classify h... ### Swin Transformer: Hierarchical Vision Transformer using Shifted Windows - Authors: Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, Baining Guo - Journal: 2021 IEEE/CVF International Conference on Computer Vision (ICCV) - Date: 2021-10-01 - DOI: https://doi.org/10.1109/iccv48922.2021.00986 - Citations: 32421 - Source: OpenAlex - llms.txt: https://science-database.com/technology/large-language-models/paper/oa-W3138516171/llms.txt - Abstract: This paper presents a new vision Transformer, called Swin Transformer, that capably serves as a general-purpose backbone for computer vision. Challenges in adapting Transformer from language to vision arise from differences between the two domains, such as large variations in the scale of visual ent... ### Exploring the Limits of Transfer Learning with a Unified Text-to-Text\n Transformer - Authors: Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J. Liu - Journal: arXiv (Cornell University) - Date: 2019-10-23 - DOI: https://doi.org/10.48550/arxiv.1910.10683 - Citations: 8346 - Source: OpenAlex - llms.txt: https://science-database.com/technology/large-language-models/paper/oa-W4288089799/llms.txt - Abstract: Transfer learning, where a model is first pre-trained on a data-rich task\nbefore being fine-tuned on a downstream task, has emerged as a powerful\ntechnique in natural language processing (NLP). The effectiveness of transfer\nlearning has given rise to a diversity of approaches, methodology, and\np... ### BioBERT: a pre-trained biomedical language representation model for biomedical text mining - Authors: Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, Jaewoo Kang - Journal: Bioinformatics - Date: 2019-09-05 - DOI: https://doi.org/10.1093/bioinformatics/btz682 - Citations: 7375 - Source: OpenAlex - llms.txt: https://science-database.com/technology/large-language-models/paper/oa-W2911489562/llms.txt - Abstract: MOTIVATION: Biomedical text mining is becoming increasingly important as the number of biomedical documents rapidly grows. With the progress in natural language processing (NLP), extracting valuable information from biomedical literature has gained popularity among researchers, and deep learning has... ### Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting - Authors: Haoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang, Jianxin Li, Hui Xiong, Wancai Zhang - Journal: Proceedings of the AAAI Conference on Artificial Intelligence - Date: 2021-05-18 - DOI: https://doi.org/10.1609/aaai.v35i12.17325 - Citations: 6866 - Source: OpenAlex - llms.txt: https://science-database.com/technology/large-language-models/paper/oa-W3177318507/llms.txt - Abstract: Many real-world applications require the prediction of long sequence time-series, such as electricity consumption planning. Long sequence time-series forecasting (LSTF) demands a high prediction capacity of the model, which is the ability to capture precise long-range dependency coupling between out... ### ChatGPT for good? On opportunities and challenges of large language models for education - Authors: Enkelejda Kasneci, Kathrin Seßler, Stefan Küchemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan Günnemann, Eyke Hüllermeier, Stephan Krusche, Gitta Kutyniok, Tilman Michaeli, Claudia Nerdel, Jürgen Pfeffer, Oleksandra Poquet, Michael Sailer, Albrecht Schmidt, Tina Seidel, Matthias Stadler, J. Weller, Jochen Kühn, Gjergji Kasneci - Journal: Learning and Individual Differences - Date: 2023-03-09 - DOI: https://doi.org/10.1016/j.lindif.2023.102274 - Citations: 6362 - Source: OpenAlex - llms.txt: https://science-database.com/technology/large-language-models/paper/oa-W4323655724/llms.txt ### Restormer: Efficient Transformer for High-Resolution Image Restoration - Authors: Syed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat, Fahad Shahbaz Khan, Ming–Hsuan Yang - Journal: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) - Date: 2022-06-01 - DOI: https://doi.org/10.1109/cvpr52688.2022.00564 - Citations: 3934 - Source: OpenAlex - llms.txt: https://science-database.com/technology/large-language-models/paper/oa-W4225672218/llms.txt - Abstract: Since convolutional neural networks (CNNs) perform well at learning generalizable image priors from large-scale data, these models have been extensively applied to image restoration and related tasks. Recently, another class of neural architectures, Transformers, have shown significant performance g... ### Large language models encode clinical knowledge - Authors: Karan Singhal, Shekoofeh Azizi, Tao Tu, S. Sara Mahdavi, Jason Lee, Hyung Won Chung, Nathan Scales, Ajay Kumar Tanwani, Heather Cole-Lewis, Stephen Pfohl, Perry W. Payne, Martin Seneviratne, Paul Gamble, Christopher Kelly, Abubakr Babiker, Nathanael Schärli, Aakanksha Chowdhery, P. Mansfield, Dina Demner‐Fushman, Blaise Agüera y Arcas, Dale R. Webster, Greg S. Corrado, Yossi Matias, Katherine Chou, Juraj Gottweis, Nenad Tomašev, Yun Liu, Alvin Rajkomar, Joëlle Barral, Christopher Semturs, Alan Karthikesalingam, Vivek Natarajan - Journal: Nature - Date: 2023-07-12 - DOI: https://doi.org/10.1038/s41586-023-06291-2 - Citations: 3744 - Source: OpenAlex - llms.txt: https://science-database.com/technology/large-language-models/paper/oa-W4384071683/llms.txt - Abstract: Abstract Large language models (LLMs) have demonstrated impressive capabilities, but the bar for clinical applications is high. Attempts to assess the clinical knowledge of models typically rely on automated evaluations based on limited benchmarks. Here, to address these limitations, we present Mult... ### Transformer-XL: Attentive Language Models beyond a Fixed-Length Context - Authors: Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V. Le, Ruslan Salakhutdinov - Date: 2019-01-01 - DOI: https://doi.org/10.18653/v1/p19-1285 - Citations: 3204 - Source: OpenAlex - llms.txt: https://science-database.com/technology/large-language-models/paper/oa-W2964110616/llms.txt - Abstract: Transformers have a potential of learning longer-term dependency, but are limited by a fixed-length context in the setting of language modeling.We propose a novel neural architecture Transformer-XL that enables learning dependency beyond a fixed length without disrupting temporal coherence.It consis... ### Point Transformer - Authors: Hengshuang Zhao, Li Jiang, Jiaya Jia, Philip H. S. Torr, Vladlen Koltun - Journal: 2021 IEEE/CVF International Conference on Computer Vision (ICCV) - Date: 2021-10-01 - DOI: https://doi.org/10.1109/iccv48922.2021.01595 - Citations: 2359 - Source: OpenAlex - llms.txt: https://science-database.com/technology/large-language-models/paper/oa-W4214755140/llms.txt - Abstract: Self-attention networks have revolutionized natural language processing and are making impressive strides in image analysis tasks such as image classification and object detection. Inspired by this success, we investigate the application of self-attention networks to 3D point cloud processing. We de... ### PaLM: Scaling Language Modeling with Pathways - Authors: Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek S. Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James T. Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier García, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, D. Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Érica Rodrigues Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Díaz, Orhan Fırat, Michele Catasta, Wei, Jason, Meier-Hellstern, Kathy, Douglas Eck, Jeff Dean, Slav Petrov, Noah Fiedel - Journal: arXiv (Cornell University) - Date: 2022-04-05 - DOI: https://doi.org/10.48550/arxiv.2204.02311 - Citations: 2136 - Source: OpenAlex - llms.txt: https://science-database.com/technology/large-language-models/paper/oa-W4224308101/llms.txt - Abstract: Large language models have been shown to achieve remarkable performance across a variety of natural language tasks using few-shot learning, which drastically reduces the number of task-specific training examples needed to adapt the model to a particular application. To further our understanding of t... ### Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding - Authors: Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J. Fleet, Mohammad Norouzi - Journal: arXiv (Cornell University) - Date: 2022-05-23 - DOI: https://doi.org/10.48550/arxiv.2205.11487 - Citations: 2113 - Source: OpenAlex - llms.txt: https://science-database.com/technology/large-language-models/paper/oa-W4281485151/llms.txt - Abstract: We present Imagen, a text-to-image diffusion model with an unprecedented degree of photorealism and a deep level of language understanding. Imagen builds on the power of large transformer language models in understanding text and hinges on the strength of diffusion models in high-fidelity image gene... ### TinyBERT: Distilling BERT for Natural Language Understanding - Authors: Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Dong Chen, Linlin Li, Fang Wang, Qun Liu - Date: 2020-01-01 - DOI: https://doi.org/10.18653/v1/2020.findings-emnlp.372 - Citations: 1707 - Source: OpenAlex - llms.txt: https://science-database.com/technology/large-language-models/paper/oa-W3105966348/llms.txt - Abstract: Language model pre-training, such as BERT, has significantly improved the performances of many natural language processing tasks. However, pre-trained language models are usually computationally expensive, so it is difficult to efficiently execute them on resourcerestricted devices. To accelerate in... ### The great Transformer: Examining the role of large language models in the political economy of AI - Authors: Dieuwertje Luitse, Wiebke Denkena - Journal: Big Data & Society - Date: 2021-07-01 - DOI: https://doi.org/10.1177/20539517211047734 - Citations: 189 - Source: OpenAlex - llms.txt: https://science-database.com/technology/large-language-models/paper/oa-W3202773593/llms.txt - Abstract: In recent years, AI research has become more and more computationally demanding. In natural language processing (NLP), this tendency is reflected in the emergence of large language models (LLMs) like GPT-3. These powerful neural network-based models can be used for a range of NLP tasks and their lan... ### Transformers and large language models in healthcare: A review - Authors: Subhash Nerella, Sabyasachi Bandyopadhyay, Jiaqing Zhang, Miguel Á. Contreras, Scott Siegel, Aysegül Bumin, Brandon Silva, Jessica Sena, Benjamin Shickel, Azra Bihorac, Kia Khezeli, Parisa Rashidi - Journal: Artificial Intelligence in Medicine - Date: 2024-06-05 - DOI: https://doi.org/10.1016/j.artmed.2024.102900 - Citations: 169 - Source: OpenAlex - llms.txt: https://science-database.com/technology/large-language-models/paper/oa-W4399367209/llms.txt --- Generated by science-database.com — The Knowledge Interface Full data available at: https://science-database.com/api/v1/technology/large-language-models